Part IV · Chapter 14
A Leadership Agenda for Open, Merit-Based Learning
The transition begins with one decision: move authority closer to valid, contestable evidence of capability.
Education reform often begins with a destination so large that no institution can act without waiting for everyone else. A reform can declare transformation, purchase a platform, and add reporting while leaving its underlying authority structure intact. That is the institutional risk this agenda seeks to address, not a claim that every reform follows one observed cycle. A practical agenda begins smaller and closer to the decision. Choose one gate, identify the legitimate purpose behind it, and rebuild the route around better evidence.
The first step is an evidence map. List consequential decisions (placement, progression, graduation, admission, hiring, licensure, access to advanced work) and the proxies each decision uses. For every proxy, state the capability being inferred, the evidence supporting that inference, the conditions under which the measure is valid, the consequence of error, and the available appeal. This exercise makes weak links visible before technology enters.
The second step is to improve one feedback loop. Select a capability important enough to matter and narrow enough to observe. Define a progression, collect representative work over time, return actionable feedback, and give learners a chance to revise and demonstrate transfer. The loop should help teachers and learners make decisions before it is used for accountability. If the system cannot improve learning at small scale, aggregation will not rescue it.
The third step is to create an alternative route. Keep the legitimate standard and allow a learner to satisfy it through direct evidence rather than only institutional history. The alternative might be a performance assessment, portfolio defense, supervised simulation, challenge examination, or validated work sample. Compare outcomes and burdens with the traditional route. An alternative that is technically available but harder to discover, finance, schedule, or appeal is not truly open.
The fourth step is governance. Establish who owns the claim, who can approve measures, who reviews bias and privacy, who hears appeals, how vendors are constrained, and when evidence expires. Use a risk process proportionate to consequence. The NIST AI RMF’s govern-map-measure-manage cycle offers a useful model for continuous oversight rather than one-time approval. 1 Public reporting should include limitations and incidents, not only success metrics.
The fifth step is to protect teacher and learner agency. Redesign workload so people have time to interpret evidence. Make system recommendations inspectable and overridable. Create protected practice records. Train participants in assessment literacy: what a measure can support, what uncertainty means, and how to identify an invalid inference. Evidence systems fail when only analysts understand them.
The sixth step is equality through clear standards and usable learning opportunities. Publish the criteria, make practice available, explain the requirements for advanced work, and provide a direct route to demonstrate readiness. The World Bank and partners’ learning-poverty work documents substantial foundational learning shortfalls. 2 My response is individual diagnosis, purposeful practice, and explicit mastery checks. Judge each learner’s work against the relevant standard rather than substitute an average or institutional label for evidence of capability.
The seventh step is disciplined scale. Expansion can change training, support, data quality, incentives, and relationships. In the teacher-coaching synthesis, larger effectiveness trials reported smaller average effects than smaller efficacy trials; this comparison raises a scale question, not a law that every intervention must deteriorate. 3 My recommendation is to expand only after specifying the intended mechanism, minimum conditions, failure signals, and resources required to preserve them.
Finally, maintain institutional humility. Some gates protect the public. Some shared sequences support coherent learning. Some credentials carry information that direct assessment cannot cheaply reproduce. The agenda is not to remove every boundary. It is to require every consequential boundary to justify itself against its purpose, evidence, burden, alternatives, and capacity for correction.
International Differences in Educational Systems
Introduction
International comparison is useful when it asks a defined question and respects the limits of the available evidence. It is much less useful when a country becomes a character in a success story: one nation represents freedom, another discipline, a third technology. Those portraits hide differences within systems and confuse description with causal explanation.
PISA offers comparative evidence about participating 15-year-olds in specified domains, with sampling and measurement qualifications. It does not experimentally assign national policies. A score difference cannot by itself establish that testing, teacher autonomy, digital tools, or a cultural trait caused the result. 4
Factors Influencing Differences in Educational Systems
A serious comparison should identify the institutions and conditions actually being compared: access to schooling, curriculum, language, teacher preparation, assessment, resources, pathways, and the population represented. The analysis should specify a jurisdiction and period rather than treating “Japan,” “the United States,” or any other country as one unchanging teaching method.
My recommendation is to begin with a problem your institution can act on. Ask which feature of another system addresses that problem, what evidence supports it, and which local conditions differ. An attractive national narrative is not a sufficient implementation plan.
Case Studies of Different Educational Systems
A case study needs an identifiable intervention, participants, period, outcome, and basis for inference. The Kenya textbook experiment discussed in Chapter 7 is one such case: its results make curricular fit and differential benefit concrete. It does not establish that every textbook investment fails or succeeds. 5
For Finland, Singapore, Germany, Kenya, or another system, the same discipline applies. Select a specific policy and inspect its evidence rather than attributing a national outcome to an appealing list of features. If a study cannot separate the policy from other changes, describe the uncertainty instead of supplying a causal story.
Impact of Differences in Educational Systems
System differences can be associated with outcome differences without identifying a transferable mechanism. School participation, prior opportunity, selection into pathways, demographic composition, and assessment coverage may all affect a comparison. An observed pattern is a starting point for inquiry, not a verdict about an entire culture.
A hypothetical leadership team might compare access to advanced mathematics across two systems. It should examine who reaches the measured stage, which alternatives remain available, and what support preceded the outcome. A high average among those selected into a route does not alone demonstrate that the route serves everyone well.
The Role of International Collaboration in Education
International collaboration can make research, materials, and professional experience available for scrutiny. Its value is the opportunity to compare explanations and learn from implementation, not a guarantee that borrowing a practice will improve outcomes.
A proposed exchange should produce a specific learning record: what the partner actually did, what it cost, who participated, what evidence changed, and what failed. Give the host institution’s account serious attention while distinguishing its interpretation from independently evaluated results. That makes collaboration an inquiry rather than an endorsement tour.
Ethical Considerations in International Educational Integration
An imported practice should meet local accessibility, privacy, curricular, and public-responsibility requirements. Consultation is needed with learners and educators who will experience the change. Translation alone does not establish that an assessment or resource carries the same meaning in another language and setting.
The recommendation is to preserve the legitimate learning goal while evaluating the local means of reaching it. A system should not gain authority simply because it is foreign, fashionable, or associated with a high-scoring jurisdiction. Equally, local familiarity should not exempt the existing approach from scrutiny.
Conclusion
International evidence can challenge assumptions without providing a national recipe. The question is not which country should be copied, but which intervention deserves a bounded local test and what would count against it.
A credible recommendation to learn across borders can stand on explicit comparison, transparent reasoning, and willingness to discover that a proposed transfer does not work. It does not require treating national reputations as experimental evidence.
Trends in International Educational Systems
Introduction
A trend is a description requiring dated evidence, not a forecast disguised as fact. Technology adoption, assessment reform, and interest in interdisciplinary learning need not move together, and adoption is not proof of benefit.
UNESCO’s 2023 technology report examines potential uses alongside uneven access, weak evidence for many products, costs, and governance concerns. 6 Leaders should therefore distinguish how widespread a practice is from whether it is effective. This section treats prominent reform ideas as choices to evaluate, not as an inevitable global transition toward merit.
Global Shift Towards Digital Learning
Digital delivery can make resources reachable beyond a particular classroom, provided learners have meaningful access. ITU’s 2025 estimates document continuing gaps in internet use and affordability; the existence of network coverage is not the same as the ability to participate effectively. 7
My recommendation is to map the actual access chain: device, connection, power, language, accessibility, time, support, and a workable alternative. A digital-first reform should not require a learner to overcome unrelated barriers merely to produce evidence of capability. Measure who cannot participate, not just activity among those who can.
Emphasis on STEAM Education
STEAM brings science, technology, engineering, arts, and mathematics into a shared design space. That description does not establish general creativity, employability, or far transfer. Each proposed course must identify the knowledge and practices it intends to develop.
A hypothetical environmental-design project could require measurement, modeling, visual communication, and argument. Assess those contributions explicitly and examine whether learners can use relevant ideas in a changed problem. Calling a project interdisciplinary does not remove the need for disciplinary depth or evidence of individual understanding.
Rise of Competency-Based Education
Competency-based progression makes an explicit capability standard central to advancement. The quality of the approach depends on the standard, the evidence elicited, scoring, support, and opportunities to try again. It is not made objective simply by removing calendar time from a progression rule.
A proposed alternative route could allow a learner to demonstrate readiness without repeating a familiar course. Check whether the route is actually available, whether its evidence is comparable in meaning, and whether unsuccessful learners receive useful feedback. A label on a curriculum is not evidence that a whole national system progresses independently of age or grade.
Increasing Importance of Lifelong Learning
The case for continued learning in this book is a normative one: people should have usable opportunities to develop capability beyond an initial period of schooling. Occupational change may create new demands, but participation in a course does not guarantee employment or advancement.
Design opportunities around the work or purpose involved. A professional might need updated practice, an accessible way to demonstrate a new skill, or feedback on a changed responsibility. The learning route should make its claim clear rather than promise a universally valuable credential.
Focus on Social and Emotional Learning (SEL) in Education
Research on school-based SEL interventions supports positive average findings across a varied literature. It does not turn one emotional-intelligence score into a general measure of merit or establish that every program improves every outcome. 89
For leadership, the proposed priority is to define the support or practice being taught, its intended benefit, and how it will be evaluated without intrusive inference. A lesson in conflict resolution is different from continuous emotion classification. The evidence for the first should not be used to authorize the second.
Impact of Globalization on Education
Cross-border study and collaboration raise questions about language, recognition, curricular differences, access, and public trust. A common credential may ease some comparisons while still leaving uncertainty about an individual’s capability. A portable artifact may offer rich evidence while requiring effort to interpret.
My recommendation is to state what a receiving institution needs to know and provide evidence close to that claim. International recognition should not depend solely on a brand, but neither should portability imply automatic equivalence. Transparent standards and a route to review remain necessary.
Conclusion
No set of fashionable trends guarantees a fairer learning system. Digital delivery, competency frameworks, interdisciplinary projects, and international exchange can each create opportunities or new barriers depending on their implementation.
A leadership agenda should therefore ask what changed for the learner: access to worthwhile work, useful support, stronger capability, or a more defensible decision. Adoption counts are not a substitute. A reform deserves continuation when its evidence and consequences justify continuation, not when its vocabulary sounds aligned with the future.
Improving Education Locally and Globally
Introduction to the Importance of Education
Education has public, personal, cultural, and occupational purposes that should be debated openly. This book’s argument is that institutions should help people develop and demonstrate capability without claiming a monopoly over every legitimate route.
That argument does not require an unsupported equation between a country’s educational spending and its economic growth. Where a specific economic or social effect is claimed, it requires its own evidence. Here the leadership task is narrower: identify an educational responsibility within reach and improve how the institution fulfills it.
Understanding the Local Education Landscape
Begin with the actual community: its learners, languages, histories, resources, obligations, and opportunities. Do not assume that a locally themed curriculum automatically produces employment or that every learner wants the same future.
A hypothetical coastal school might use a marine-science project to investigate water quality. It should still teach the relevant scientific concepts, provide accessible participation, and evaluate the quality of explanations. Local relevance is a design rationale, not proof of a measured learning gain or a mandate to confine learners to local occupations.
Identifying Challenges in Local Education
A problem statement should distinguish observed conditions from inferred causes. Missing assignments could reflect misunderstanding, inaccessible materials, work obligations, or a reporting error. A low score can identify a concern without showing which intervention will help.
Ask learners and educators what prevented the intended work, inspect representative evidence, and compare competing explanations. The outcome may call for better teaching, a schedule change, an accommodation, additional resources, or a repair to the measure itself. Starting with an explanation that only the preferred platform can solve prejudges the inquiry.
Strategies for Improving Local Education
My recommendation is to choose a local problem before choosing a platform: a bottleneck in feedback, an inaccessible resource, or an administrative demand that displaces teaching. Analytics and adaptive systems are possible responses, not established solutions because they are innovative. 6
A hypothetical school could compare a bounded adaptive-practice pilot with its current approach, assessing later unaided work, access, workload, and cost. The pilot should have a stopping rule. Increased platform activity alone does not justify expansion, and a negative result should be useful evidence rather than a reputational threat.
Understanding the Global Education Landscape
The World Bank and partners’ 2022 learning-poverty estimate described a severe foundational-learning problem using specific definitions and pandemic-era modeling. It is not a continuously updated count of every child or a direct measure of all educational purposes. 2
Use large-scale indicators to identify questions and responsibilities, then obtain relevant local evidence. A national or global average cannot diagnose an individual learner. Nor does an indicator identify the cause of a shortfall simply because the problem is urgent.
Strategies for Improving Global Education
Sharing research is valuable when the conditions travel with the finding. Record who participated, which language and curriculum were used, the resources supplied, what the comparison received, and the duration of follow-up.
The proposed discipline is to preserve those conditions in any adaptation plan. Identify what can be retained, what must change, and which claim will require new evidence. A program tested under substantial support should not be represented as a low-support solution merely because its software can be copied cheaply.
The Role of Technology in Improving Education
Technology may help with a defined instructional or administrative task. The lesson-preparation trial discussed in Chapter 5 provides evidence about preparation time in a bounded teaching context, not automatic grading reliability or improved student attainment. 10
A leader should name the benefit sought and count the work required to obtain it. Time spent checking generated material, supporting access, correcting errors, and handling appeals belongs in the evaluation. Faster output matters only in relation to the quality and purpose of that output.
Conclusion of Local and Global Education Improvement
The case for improvement should remain open to alternatives. A teacher-led approach may be better for one problem; a digital resource may help with another. A governance repair may be more important than either.
Judge proposals by evidence, access, sustainability, and consequences. My recommendation is to publish the reasoning behind local choices so others can learn from them without mistaking one favorable outcome for a universal recipe. An institution contributes to shared learning when it makes its uncertainties and failures usable as well as its successes.
The Possibility of Improving Education Globally
Introduction to Global Educational Advancement
Possibility is not inevitability. The evidence assembled in this book supports several bounded practices and also documents failures, mixed findings, and weak inferences. A credible reform agenda uses both.
The leadership commitment is to make improvement testable: identify the present problem, a plausible response, a comparison, and a consequence worth observing. Neither institutional tradition nor technological optimism should be allowed to claim success without that discipline.
Challenges in Global Education
Access, instructional quality, resources, language, safety, and public responsibility present different problems. An intervention directed at one may leave another unchanged. For example, distributing material does not establish that the material is appropriate to a learner’s prior knowledge.
The Kenya textbook experiment illustrates why the match between a resource and its users matters. 5 The lesson is not to abandon provision, but to investigate who can use what is provided and whether additional support or different materials are needed.
The Potential of Technology in Global Education
Remote access can remove a geographic barrier while leaving other barriers intact. A lecture delivered online still requires usable access, relevant language, time, and an opportunity to ask for help. UNESCO’s technology report examines educational effectiveness and sustainability. 6 My recommendation is to test whether the delivery arrangement supports useful practice and valid demonstrations of mastery before extending it.
My recommendation is to evaluate the whole access pathway. Record learners who were unable to begin or continue, not only successful participants. A platform does not create equal opportunity merely by making a registration page publicly available.
The Role of Personalized Learning in Global Education
Personalization should be treated as a set of testable decisions about support. Define what is adapted and why. Evidence from a successful package of adaptive practice and human support does not establish that the same software will work without that support.
The India study discussed in Chapter 11 is a useful example of a bounded positive result. 11 A new implementation should document which features and conditions differ. That is how a promising intervention becomes a researchable local proposal rather than a claim that the problem has already been solved.
The Importance of Teacher Training and Support
Teacher development needs a clear instructional purpose and time to change practice. Coaching research reports positive average effects while raising questions about maintaining them at scale. 3 That is more informative than asserting that any training session will improve student outcomes.
For a proposed assessment-literacy program, ask teachers to examine real work, compare interpretations, test feedback, and revisit decisions. Evaluate the practice that changed, not just attendance at the session. Support should respond to implementation difficulties rather than treating nonadoption as evidence of resistance.
The Potential of Collaborative Learning
Collaboration creates an opportunity for explanation, disagreement, joint production, and feedback. It can also conceal unequal participation or leave an individual learner’s understanding unclear. A recent synthesis of technology-supported collaborative interventions reports positive average outcomes, but it does not validate every shared-document feature or group assignment. 12
A proposed cross-school project should make roles, access, individual contributions, and the learning goal explicit. Assess both the joint product and what each learner can explain. The opportunity to collaborate and the demonstrated benefit of collaboration are distinct claims.
The Role of Assessment in Global Education
Assessment should follow the claim and the intended use. A shared item format may aid comparison while still requiring evidence about translation, curricular coverage, accessibility, and the population represented.
My recommendation is to preserve direct work behind summaries and to explain the limits of cross-context comparison. A higher score on a particular assessment is not a complete ranking of people, cultures, or education systems. The decision must remain proportionate to what the evidence actually samples.
The Future of Global Education
A future forecast cannot substitute for present evidence. AI, adaptive content, and new credentials may change available options, but institutions still decide who may participate, what standards matter, and how errors can be challenged.
Build a process that can respond to improvement and deterioration alike. Preserve alternatives, record versions and conditions, and reevaluate when the population or purpose changes. A system deserves continuing authority through the quality of its evidence and accountability, not because it was once introduced as the future.
Conclusion of Global Education Improvement
The proposal is a disciplined form of openness: accept useful evidence from multiple routes, retain legitimate public standards, make opportunity visible, and allow consequential decisions to be challenged. Those are policy commitments informed by research, not outcomes already established by adopting a slogan.
The following validation ladder turns that commitment into a sequence of decisions. Its purpose is not to demand certainty before any experiment. It is to prevent a promising idea from acquiring greater authority than its evidence can support.
A validation ladder before scale
Innovation should move through increasingly demanding forms of evidence. Begin with technical and content checks: does the system function, cite correctly, preserve accessibility, and avoid known prohibited behavior? Continue with usability and bounded learning studies. Then use partner pilots, credible comparison groups, independent evaluation, and longitudinal evidence when the decision and scale warrant them. A successful demonstration is not a school-wide effectiveness study.
Each stage should have a predeclared learning claim, comparison, population, duration, outcome, failure threshold, check for consistent assessment standards, and stopping rule. Engagement and satisfaction can explain adoption but cannot substitute for retention, transfer, or independent performance. When a tool changes rapidly, version, prompts, model, content, and human support belong in the evidence record.
The U.S. Department of Education’s AI guidance treats teachers and other people as central to the design and operation of educational systems, not as ceremonial reviewers. 13 UNESCO’s guidance also addresses human agency, privacy, age appropriateness, transparency, and validation across the life cycle. 14 I recommend making these safeguards operational requirements. A vendor’s claim to follow them should be translated into testable contract terms, observable workflow, learner rights, and named institutional accountability.
For high-risk decisions, leaders should require independent evidence and an appeal route before deployment. For low-risk assistance, they can permit narrower experimentation while logging failure and preserving alternatives. The purpose of a validation ladder is not to prevent learning organizations from trying anything. It is to keep the strength of a decision proportional to the strength of the evidence.
Build the conditions for an honest pilot
A pilot’s interpretation is weakened when comparison conditions are starved, temporary specialist support is omitted from the proposed operating model, or success is defined only after results are known. Coerced participation also creates an ethical problem regardless of whether a statistical comparison remains possible. Before launch, describe ordinary practice honestly and resource it normally. Record training time, troubleshooting, added staffing, licensing, devices, connectivity, accommodations, and work transferred to families. Those costs are part of the intervention.
Select participants in a way that makes the intended inference possible. A volunteer pilot may be appropriate for usability and early failure detection, but it cannot establish that the system will work for everyone. Test the actual reading levels, access requirements, schedules, equipment, and instructional conditions the proposed service must support. Define the same performance criteria before examining results. This is a test of operational readiness, not a substitute for evidence that learners can perform the intended work.
Protect the comparison. Teams naturally want a promising pilot to succeed and may provide it with better staffing, smaller groups, or more attention. Document those differences. If the new model requires them, they are not contamination; they are part of the cost. If they cannot be sustained, the pilot tested a temporary program rather than the proposed operating model.
Predeclare what would count as benefit, harm, uncertainty, and failure. Include retention and transfer when the claim requires them, not only completion and satisfaction. Examine work samples and implementation records alongside numerical outcomes. Invite participants to report harms that the metric did not anticipate. A result can be statistically positive and operationally unacceptable if it depends on surveillance, inaccessible routines, or unsustainable labor.
At the decision point, choose among more than scale or abandon. Continue the pilot to reduce uncertainty; narrow use to a population or function with stronger evidence; change the workflow; remove a risky data source; require human review; or stop. Publish the reasoning internally so later leaders know what was learned rather than inheriting a slogan about success.
Scale in stages and keep a holdout or other credible monitoring strategy where feasible. Early adopters are often unrepresentative, and system behavior changes with volume. Track override, appeal, differential error, staff workload, learner dependence, and opportunity effects after launch. A procurement renewal is another evidence decision, not an administrative default.
Leadership is most visible when evidence is inconvenient. A credible leader protects the person who reports failure, tells a funder that an attractive result is too narrow, and stops a system whose sunk cost has become its strongest argument. Open, merit-based learning requires institutions to prove their own claims with the same seriousness they demand from learners.
Reading a pilot result without turning it into a slogan
Consider a hypothetical proposal to replace part of a course with adaptive practice. The following questions are an author’s decision exercise, not an account of a successful school. They show how to preserve the difference between a promising program and a warranted institutional decision.
First, identify the comparison. If the new program adds an hour of supervised practice while the comparison receives no additional time, a favorable outcome concerns the added program. It does not isolate the contribution of adaptation. That may still be an important finding: the institution may be deciding whether to fund the entire additional program. The problem is not that the package contains several elements; it is reporting a package effect as proof of one element.
Second, examine the assessment condition. If learners can use the tutor while producing the final answer, the result concerns supported performance. That is relevant when the intended capability includes competent tool use. It is insufficient when the decision requires independent execution. A separate no-assistance component should be justified by the responsibility involved, not by a general suspicion of tools. Delayed and changed tasks supply further evidence when durability or transfer is part of the claim.
Third, inspect missing participation and missing outcomes. A favorable average among completers cannot automatically represent people who could not access the program or left it. Record why observations are missing and how the analysis treats them. Do not quietly remove difficult experiences from the denominator. Equally, do not assume every missing record represents failure; uncertainty should remain visible until there is evidence to interpret it.
Fourth, distinguish statistical uncertainty from practical importance. A result too imprecise to distinguish alternatives does not establish that the alternatives are equivalent. A detectable average difference may be too small to justify the cost or burden. The decision needs a predeclared account of what improvement matters and what harm is unacceptable. That account includes educational values as well as statistical reasoning.
Fifth, count the operating model. Preparation, review, technical support, devices, subscriptions, accommodations, and learner time belong in the comparison. Work transferred to families does not disappear because it is absent from a school budget. If the pilot relied on unusually intensive attention, record that attention as part of what was tested. A cheaper scaled version is a changed intervention and should not inherit the original result without scrutiny.
Sixth, decide how much authority the result can support. A short voluntary pilot may justify another low-risk trial while remaining inadequate for compulsory adoption or consequential placement. A positive average may support a narrower use than the procurement proposal. An institution should be able to continue, revise, restrict, or stop without treating every cautious decision as opposition to innovation.
This reading discipline does not weaken favorable evidence. It gives that evidence a clear job. A credible positive finding becomes a reason to test or adopt a defined practice under defensible conditions. A limitation becomes a design requirement or an unanswered question, not an excuse to ignore the entire result. The aim is an organization that can act on evidence without converting uncertainty into either paralysis or exaggerated certainty.
A ninety-day start
In the first thirty days, inventory gates and choose one bounded decision. In the next thirty, convene learners, educators, decision owners, accessibility and privacy expertise, and people currently excluded by the route. Define the capability claim and design a direct-evidence alternative. In the final thirty, test the feedback loop with a small voluntary group, document failures, and publish what will be changed before any consequence attaches.
Do not begin with a universal learner record, predictive platform, or new credential taxonomy. Begin with a valid claim, real work, and a decision someone is willing to revise. Infrastructure should follow proven need.
Implications for leaders
- Inventory consequential gates and the claims their proxies are expected to support.
- Improve learning feedback before attaching accountability.
- Pilot one discoverable, affordable alternative route to the same legitimate standard.
- Establish governance, appeals, retention, and vendor limits before scale.
- Scale mechanisms and conditions, not branding and user counts.
Questions to carry forward
- Which gate could be redesigned within your present authority?
- What legitimate public or educational purpose must remain protected?
- What evidence would persuade you to revise the new system or abandon it?
Notes
- 1Tabassi, Elham. 2023. “Artificial Intelligence Risk Management Framework (AI RMF 1.0)”. National Institute of Standards and Technology.↩
- 2World Bank, UNICEF, FCDO, et al.. 2022. “The State of Global Learning Poverty: 2022 Update”. World Bank.↩
- 3Kraft, Matthew A., Blazar, David, Hogan, Dylan. 2018. “The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence”. Review of Educational Research, vol. 88, no. 4, 547–588.↩
- 4OECD. 2023. “PISA 2022 Results (Volume I)”. OECD Publishing.↩
- 5Glewwe, Paul, Kremer, Michael, Moulin, Sylvie. 2009. “Many Children Left Behind? Textbooks and Test Scores in Kenya”. American Economic Journal: Applied Economics, vol. 1, no. 1, 112-135.↩
- 6Global Education Monitoring Report Team. 2023. “Global Education Monitoring Report 2023: Technology in Education: A Tool on Whose Terms?”. UNESCO.↩
- 7International Telecommunication Union. 2025. “Measuring Digital Development: Facts and Figures 2025”. ITU.↩
- 8Durlak, Joseph A., Weissberg, Roger P., Dymnicki, Allison B., et al.. 2011. “The Impact of Enhancing Students’ Social and Emotional Learning: A Meta-Analysis of School-Based Universal Interventions”. Child Development, vol. 82, no. 1, 405–432.↩
- 9Cipriano, Christina, Strambler, Michael J, Naples, Lauren H, et al.. 2023. “The state of evidence for social and emotional learning: A contemporary meta-analysis of universal school-based SEL interventions”. Child Development, vol. 94, no. 5, 1181-1204.↩
- 10Roy, Palak, Poet, Helen, Staunton, Ruth, et al.. 2024. “ChatGPT in Lesson Preparation: A Teacher Choices Trial”. Education Endowment Foundation, Evaluation report.↩
- 11Muralidharan, Karthik, Singh, Abhijeet, Ganimian, Alejandro J.. 2019. “Disrupting Education? Experimental Evidence on Technology-Aided Instruction in India”. American Economic Review, vol. 109, no. 4, 1426-1460.↩
- 12Xu, Enwei, Feng, Xuezhen, Ning, Kewei, et al.. 2025. “The effectiveness of technical-supported collaboration in promoting students’ learning outcomes: a meta-analysis based on empirical literature”. Humanities and Social Sciences Communications, vol. 12, no. 1, 1505.↩
- 13U.S. Department of Education, Office of Educational Technology. 2023. “Artificial Intelligence and the Future of Teaching and Learning: Insights and Recommendations”. U.S. Department of Education.↩
- 14Miao, Fengchun and Holmes, Wayne. 2023. “Guidance for Generative AI in Education and Research”. UNESCO.↩