Conclusion
Merit, Effort, and Mastery
The standard belongs to the work and the public purpose, not to the prestige of the path taken to reach it.
The most durable gate is the belief that capability becomes real only after an institution recognizes it. That belief confuses recognition with existence. People learn before a course begins, between formal lessons, in work and family life, through failure, through practice, and long after a credential is awarded. Institutions do not create all of that learning. At their best, they make it more likely, more visible, more connected, and more trustworthy.
The task is not to choose between institutions and freedom. It is to build institutions worthy of free learners. Such institutions state what evidence they need and why. They offer more than one defensible way to produce it. They distinguish present readiness from future potential. They make consequential judgments reviewable. They reduce the advantage of knowing the hidden rules. They treat measurement as a conversation with reality rather than a verdict on a person.
Four commitments follow from the argument of this book.
First, begin with the learning claim. Before selecting a test, platform, curriculum, credential, or dashboard, say what the learner should be able to understand or do. Then ask what evidence would warrant that conclusion and what evidence would challenge it. The sequence matters. Institutions too often begin with the data they already possess and allow available measures to define the goal.
Second, separate feedback from judgment whenever possible. Evidence collected to help someone improve can be frequent, provisional, and diagnostic. Evidence used to grant independent responsibility must meet a different standard. When every attempt becomes part of a permanent ranking, the design risks making experimentation costly and feedback harder to use openly. A healthy system creates protected space for practice and explicit moments for consequential demonstration.
Third, make the route contestable. Every consequential measure is incomplete. Learners need a way to correct errors, add context, present alternative evidence, and ask for human review. Contestability is not an administrative nuisance; it is how an evidence system admits that it can be wrong. NIST’s risk frameworks emphasize governance, documentation, measurement, and continuing management rather than one-time technical certification. 12 Education systems deserve the same discipline.
Fourth, make the route to mastery usable. Publish the requirements, provide purposeful practice, and make feedback available before a consequential demonstration. A digital route still needs workable tools and connectivity. Nearly universal mobile coverage does not mean universal meaningful access; affordability and quality gaps remain, and billions of people were still offline in the ITU’s 2025 estimates. 3 My recommendation is to inspect whether the proposed route actually works, then assess each learner against the same relevant standard.
Merit, effort, and mastery therefore describe distinct responsibilities. The learner undertakes the work. The educator supplies instruction, challenge, and feedback. The institution states and consistently applies the standard. Equality concerns the rules and the opportunity to demonstrate capability; it does not require identical results or excuse an unfulfilled safety requirement. My argument is that recognition should follow warranted evidence of mastery, while practice remains a place to improve, correct mistakes, and make the next serious attempt.
There will still be standards. A bridge must stand. A clinician must recognize danger. A teacher must be able to create a safe and intellectually serious environment. A learner who has not yet demonstrated readiness may need more practice before taking on independent responsibility. The difference is that the standard belongs to the work and the public purpose, not to the prestige of the path taken to reach it.
There will still be teachers. My argument is that abundant content does not remove the need for human attention, judgment, care, challenge, and example. The teacher’s role grows rather than shrinks when the institution stops asking the teacher to function mainly as a broadcaster, compliance monitor, and scoring machine. Coaching research suggests that sustained, context-specific support can improve instructional practice, even while results become harder to maintain at scale. 4 That is a reminder that relationships and implementation are not inefficiencies to be engineered away. They are part of the mechanism.
There will still be institutions. Their legitimacy should rest less on controlling entry and more on the quality of the environments, communities, evidence, and public trust they create. A university may remain an extraordinary place to learn without being the only place where advanced learning can become credible. A school may coordinate a rich civic childhood without pretending that a calendar age and seat-time schedule fully describe development. A professional body may protect the public while recognizing multiple pathways to the same demanding standard.
And there will still be uncertainty. Evidence does not eliminate judgment. It disciplines judgment. It reveals when a leader is making an inference, when a measure is indirect, when a sample is narrow, when an intervention depends on context, and when values, not data, are deciding the goal. The honest system does not hide those choices behind a score.
The possibility of improving education rests in that honesty. Leaders do not need to wait for a perfect technology or a national redesign. They can reduce one unnecessary gate, replace one weak proxy, protect one space for low-stakes practice, give one learner another way to demonstrate mastery, return one data point to the person it describes, and create one review process that can admit error. Those changes accumulate into a different institutional character.
Learning has never waited for permission. Our systems should stop pretending that it must.
A final set of questions for leaders
- Which decisions in your institution rely on proxies furthest from the capability being claimed?
- Where can a learner present direct evidence, and where are they required to present only institutional history?
- Which measures help learners improve, and which exist mainly because they are easy to aggregate?
- What data would you refuse to collect even if it improved prediction?
- Who can challenge a decision, and what happens when the system is wrong?
- Which alternative route could meet the same legitimate standard without reproducing the same gate?
- What would become possible if teachers spent less time administering evidence and more time interpreting it with learners?
The safeguards are part of the promise
An evidence system can repeat the mistakes of the arrangement it replaces. A polished portfolio can conceal weak reasoning. Alternative credentials can multiply until learners carry the cost of proving themselves again to every institution. Continuous assessment can turn practice into permanent surveillance. Personalized pathways can quietly lower expectations. Automated recommendations can make yesterday’s pattern into tomorrow’s ceiling. Calling these arrangements merit-based does not validate them. Their judgments must still follow demonstrated capability under clear, consistent standards.
The safeguards are therefore not secondary compliance work. They define whether the reform is actually open.
First, evidence must remain contestable. A learner needs access to the evidence used, an explanation of the inference, and a practical route to add context or demonstrate the capability again. Review cannot be so slow, expensive, or obscure that only well-resourced people can use it. When a decision has serious consequences, a human being with authority must be responsible for it.
Second, practice must remain protected. Learners need places to be confused, wrong, experimental, and unfinished without producing a permanent reputational record. A system that observes everything does not necessarily understand more. It can make honest risk-taking less likely and reward learners who know how to manage appearances. Institutions should specify when evidence is formative, when it becomes consequential, and how long each record survives.
Third, pathways must preserve ambition. Adaptation should change support, pace, representation, or sequence in response to evidence; it should not assign a smaller future because early data predicted difficulty. The learner must be able to see and challenge a recommendation, enter a more demanding route, and produce new evidence that changes the model. Personalization without upward mobility is tracking with friendlier language.
Fourth, opportunity belongs in the validity argument. If an assessment requires broadband, a quiet room, sophisticated equipment, fluent academic English, unpaid time, or insider coaching unrelated to the capability, those conditions affect what the result means. Removing irrelevant barriers does not weaken a standard. It makes the evidence closer to the intended claim.
Fifth, portability must include restraint. A learner should be able to carry relevant evidence without carrying an exhaustive history. Selective disclosure, expiration, clear provenance, and purpose limits matter more than a universal record. The person should not have to trade permanent visibility for recognition.
What an institution can become
Once recognition is separated from monopoly, institutions do not disappear. Their most valuable work becomes easier to see. They can assemble communities that sustain attention and ambition. They can provide laboratories, studios, clinics, tools, archives, and protected time. They can connect novices with expert judgment. They can curate coherent sequences without claiming those sequences are the only possible ones. They can develop assessment expertise, preserve public standards, and make difficult evidence trustworthy across contexts.
Their authority then rests on service and integrity rather than scarcity alone. A strong institution attracts learners because its environment helps them become capable, not because it can threaten to keep everyone else invisible. Its credential has value because the underlying claims, evidence, criteria, and quality controls deserve trust, not because competing evidence is prohibited.
For leaders, this reframes accountability. The institution should report more than completions and average scores. It should be able to show what learners attempted, how their work improved, where transfer was demonstrated, whether the same relevant criteria were applied, which measures failed, which decisions were appealed, and what changed as a result. Transparency should include uncertainty and failure, not merely polished outcome dashboards.
This is demanding work. It requires judgment that cannot be purchased as a platform, relationships that cannot be reduced to data exhaust, and governance that remains active after launch. It also offers a more honorable role for educational leadership: not guarding the only authorized door, but building environments in which more people can develop real capability and more valid doors can recognize it.
Learning without gatekeepers is not learning without teachers, institutions, standards, or responsibility. It is learning without arbitrary monopoly over who may try, who may be seen, and which path is allowed to count. The work begins wherever a leader can bring a decision closer to evidence and make that evidence answerable to the learner whose future it affects.
A compact operating doctrine
The argument can be reduced to an operating doctrine for institutions that want to act before every policy, standard, and technology is settled.
Name the claim before choosing the measure
Begin every consequential design with a sentence that a learner could understand: “We need evidence that you can…” Complete it with an observable capability, the conditions that matter, and the reason the claim is legitimate. If the sentence ends with attendance, possession of a credential, use of a platform, or completion of a sequence, ask what those facts are expected to represent. Sometimes the proxy is administratively necessary. It should still be named as a proxy.
Then choose evidence close enough to the claim. Foundational knowledge may require efficient sampling and delayed retrieval. Complex performance may require artifacts, observation, explanation, and repeated judgment. Transfer requires changed contexts. Responsible independence may require performance without certain forms of assistance. The measure follows the claim; the available database does not define it.
Preserve practice as a different kind of space
Learners need frequent evidence to improve and carefully bounded evidence for consequential judgment. Confusing the two corrupts both. When every practice record contributes to ranking, the system risks rewarding concealment and easy success instead of open examination of uncertainty. When consequential decisions rely on casual practice traces, institutions make broad claims from evidence never designed to carry them.
Create an explicit boundary. Formative evidence is provisional, accessible to the learner, and retained only as long as it helps action. Consequential evidence is collected under stated conditions, reviewed against visible criteria, and accompanied by an appeal route. A learner should know when that boundary is crossed.
Make every gate justify its burden
Some boundaries protect safety, coherent progression, scarce resources, or public trust. Others persist because an institution inherited them. For each gate, document its purpose, the evidence that it serves that purpose, the people it wrongly excludes, the people it may wrongly approve, and the alternative routes available.
An alternative route should not be ceremonial. It must be discoverable, affordable, accessible, and scheduled often enough to use. It must assess the same legitimate capability without adding unrelated burdens. Its evidence should receive the same serious review as evidence from the conventional path. An institution has not opened a gate if the alternative exists only in policy language.
Keep uncertainty attached to the decision
Every educational inference has limits. Samples are incomplete, contexts change, raters disagree, models drift, and people grow. Institutions often strip uncertainty away as evidence moves upward: a varied body of work becomes a rubric score, the score becomes a category, the category becomes a placement, and the placement becomes part of identity.
Keep the chain visible. Record the evidence, criteria, conditions, and degree of confidence appropriate to the decision. Do not attach clinical, psychological, or moral meaning that the evidence was not designed to support. Set a review date when capability can change. A placement should describe a present instructional decision, not a permanent kind of person.
Return evidence to the learner
Data collected in the learner’s name should create value for the learner. That can mean timely feedback, a clearer progression, a portable work sample, an explanation of a decision, or a route to challenge an error. A system that extracts detailed evidence for institutional reporting and returns only a grade is structurally misaligned.
Returning evidence can also support governance. When learners and educators can inspect the record, they can identify inaccurate context, missing work, inaccessible tasks, and inferences that do not match experience. Transparency does not eliminate professional judgment. The purpose is to make its grounds available for review and correction.
Refuse the data that would change the institution for the worse
Not every predictive advantage deserves to be used. A signal may improve classification while normalizing surveillance, chilling experimentation, or creating a record too dangerous to retain. Leaders should identify red lines before a vendor or crisis makes the choice seem inevitable.
Those lines may include biometric emotion inference, covert monitoring, sale or advertising use, indefinite practice histories, model training without a separate lawful basis, automated high-consequence decisions without effective human review, and risk scores whose inputs or correction process cannot be explained. The exact policy depends on context and law. The principle is broader: accuracy is not the only public value.
Scale trust, not merely throughput
A pilot can depend on extraordinary people, attention, and resources. Scale changes the mechanism. Expansion can introduce more varied contexts, thinner support, harder calibration, greater privacy exposure, and pressure to standardize; these risks should be evaluated rather than assumed to occur identically everywhere. Leaders should define what must remain true as participation grows: time for feedback, quality of evidence, accessibility, appeal capacity, data limits, and learner agency.
If those conditions cannot survive, the design is not ready to scale. Reducing scope may preserve more value than expanding a weakened version. The relevant success measure is not adoption. It is whether the institution continues to produce valid learning opportunities and trustworthy evidence for people unlike the original participants.
The invitation
The invitation of this book is practical. Choose one decision that currently depends on institutional history more than demonstrated capability. Name the legitimate standard. Invite the people affected to describe the hidden burdens. Build one direct-evidence route. Protect practice. Collect only what the decision requires. Make the criteria public. Give the learner the evidence and an appeal. Study who can use the route, who cannot, and why. Revise before expanding.
That work will not end gatekeeping everywhere. Its purpose is to change the character of authority where the leader actually has responsibility: an institution more capable of recognizing learning it did not authorize, more honest about what its measures cannot see, and more useful to people poorly represented by its proxies. Whether a particular reform achieves those aims remains a question for evidence.
The future of education should not depend on abolishing every institution or trusting every new technology. It should depend on a simpler and more demanding promise: wherever learning occurs, people deserve a fair chance to develop it, show it, understand how it was judged, and carry credible evidence into the next opportunity.
That promise also changes how success is recognized. The strongest evidence system is not the one that produces the most precise hierarchy. It is the one that helps more people see what they know, identify what they need, improve through serious work, and enter responsibilities for which they are genuinely prepared. It notices growth without confusing growth with readiness. It protects standards without protecting inherited monopolies. It values expert judgment without making expertise unanswerable. It uses technology without surrendering purpose to the technology’s available signals. Above all, it leaves room for a person to become different from the record accumulated about them. Education is justified by that possibility of development. Any measurement worthy of education must preserve it.
Notes
- 1Tabassi, Elham. 2023. “Artificial Intelligence Risk Management Framework (AI RMF 1.0)”. National Institute of Standards and Technology.↩
- 2National Institute of Standards and Technology. 2020. “NIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0”. NIST.↩
- 3International Telecommunication Union. 2025. “Measuring Digital Development: Facts and Figures 2025”. ITU.↩
- 4Kraft, Matthew A., Blazar, David, Hogan, Dylan. 2018. “The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence”. Review of Educational Research, vol. 88, no. 4, 547–588.↩