Part III · Chapter 9
Measuring Growth, Mastery, and Transfer
Measure close to the claim, across time, and in forms that help the learner act.
Measurement earns its place in learning when it changes what someone can do next. That is a different standard from producing a defensible rank. A useful learning measure locates present understanding, makes progress visible, identifies a productive next challenge, and returns information while action is still possible. Its first audience is the learner and the people supporting the learner.
This chapter proposes three distinct claims: growth, mastery, and transfer. Growth asks how performance changed over a defined interval under described conditions. Mastery asks whether a learner can perform a specified body of work to an acceptable standard. Transfer asks whether knowledge or skill can be selected and used when surface cues, context, time, or purpose change. These claims overlap, but no single score answers all three.
Growth is relational. It requires at least two observations and a scale on which change is meaningful. A learner can make substantial growth while remaining below a readiness threshold; another can meet the threshold with little recent growth because of earlier opportunity. Systems that collapse both into one rank lose information leaders need. The distinction also changes the moral narrative. Present performance matters, but it should not erase the distance someone has traveled or the support needed to continue.
Mastery requires a domain definition. “Understands algebra” is too broad until the institution describes concepts, representations, procedures, applications, and expected independence. Sampling matters because no assessment can include the whole domain. A mastery decision becomes stronger when evidence includes different tasks, explanations, and repeated performance rather than one familiar format. It also needs an expiration theory: some capabilities remain stable; others decay without use or change as tools and standards evolve.
Transfer is the most demanding claim and often the most casually made. A learner may remember information without recognizing when it applies. A skill can travel across examples within a classroom yet fail across disciplines, workplaces, or social settings. Transfer assessment therefore changes conditions deliberately and asks the learner to select a response rather than merely execute a cued one. 1 The institution should name how far it expects the learning to travel.
Formative assessment supports this direction when evidence changes teaching and learning. Black and Wiliam’s influential review connected classroom evidence and feedback with substantial potential gains. 2 A quiz is not formative because it is short. It is formative when the resulting information changes what the learner or teacher does while learning is still underway.
Feedback quality matters. Research syntheses distinguish information about the task, process, and self-regulation from praise or judgment directed at the self. 3 “Good job” may support a relationship but gives little guidance. “Your evidence supports the first claim but not the second; compare the counterexample and revise your rule” creates a possible next action. Measurement without an action pathway is reporting, not learning support.
Leaders need evidence portfolios at two levels. At the learner level, evidence should show claims, artifacts, observations, context, feedback, revisions, and current status. At the system level, leaders need patterns across groups, opportunities, tasks, and time without pretending that the aggregate explains every individual. The link between levels must remain inspectable: a dashboard should lead back to real work, not become a world sealed off from it.
Validity begins with the decision
Measurement is often introduced through instruments: tests, rubrics, dashboards, portfolios, observations, or analytics. Leaders should begin one step later, with the decision. What will someone do because this evidence exists? The same measure may be reasonable for choosing the next practice problem and irresponsible for denying graduation. Consequence changes the burden of evidence.
A useful framework separates five questions. Relevance: does the task elicit the capability the decision claims to concern? Coverage: does the evidence represent enough of the domain? Generalization: would performance persist across time, tasks, settings, or evaluators? Fairness: do irrelevant barriers change what the evidence means for different learners? Consequence: what harms follow from false confidence, false rejection, delay, or misclassification?
No measure answers all five by itself. A multiple-choice item can sample knowledge efficiently while missing explanation and performance. A rich project can reveal planning, integration, and revision while sampling a narrow topic and making outside assistance hard to interpret. A supervisor’s observation can capture real conditions while introducing local expectations and interpersonal bias. Multiple measures help only when their weaknesses differ and when leaders know how the evidence will be combined.
Growth is not a lower standard
Growth and mastery answer different questions. Mastery asks whether present performance meets a defined level. Growth asks how capability changed over time. A learner can make extraordinary growth and still need more practice before assuming a safety-critical responsibility. Another can meet a threshold with little recorded growth because prior opportunities placed them near it at the beginning. Both facts matter.
Institutions get into trouble when they force one score to carry both meanings. Current status answers whether the learner has met a stated requirement; growth answers how performance has changed through learning and practice. Neither should impersonate the other. An evidence profile should preserve both, along with the conditions under which the evidence was produced, so effort can receive useful feedback while readiness remains a separate judgment.
Growth measures also require caution. Different starting points, ceiling effects, regression toward the mean, test scaling, missing data, and changes in instruction can affect interpretations. A precise number may imply more comparability than the design warrants. Leaders should treat growth as an estimate with assumptions, not a property residing inside a learner.
Mastery needs a visible claim
The word mastery is frequently used without a domain, condition, or horizon. Mastery of what, to what standard, with which tools, after how much delay, and with how much independence? A learner may execute a procedure fluently and lack conceptual explanation; reason well with support and fail under time pressure; perform in a simulation and struggle amid real social complexity. These are not reasons to abandon mastery. They are reasons to define it.
A defensible mastery claim names essential knowledge and performance, acceptable variation, common failure modes, and the evidence that would count against readiness. It distinguishes assistance that belongs to competent practice from assistance that invalidates the claim. Professionals use references, instruments, checklists, colleagues, and software. An assessment that prohibits every authentic tool may measure memory under artificial constraint rather than competent work.
Transfer must be designed and tested
Transfer is not a bonus automatically produced by understanding. It depends on what the learner notices, how knowledge was represented, the similarity and difference between contexts, prior experience, and opportunities to practice selecting a strategy. Frameworks for transfer emphasize dimensions such as domain, physical and social context, time, function, and modality. 1
Instruction can support transfer by varying surface features while preserving deep structure, comparing cases, asking learners to explain when a method applies, and requiring them to choose among strategies. Assessment must then change the prompt. Repeating a rehearsed format may establish fluency but cannot, by itself, establish far transfer.
Leaders should be suspicious of broad claims based on narrow evidence. Improvement on a rewarded test does not necessarily represent improvement across the wider domain, a concern emphasized in reviews of test-based accountability. 4 The appropriate response is not to dismiss standardized evidence wholesale. It is to match the breadth of the conclusion to the breadth and independence of the evidence.
An evidence portfolio, not a data pile
An evidence portfolio is a structured argument. It should not be an unlimited archive of learner activity. For each capability, it selects representative evidence, records conditions and assistance, applies visible criteria, and preserves feedback and revision when growth matters. It can include work samples, explanations, demonstrations, simulations, observations, and appropriately designed assessments.
Quality matters more than volume. More records can create the appearance of confidence while repeating the same bias or sampling the same narrow condition. Evidence should be retained because it changes a decision, supports feedback, documents progress, or satisfies a legitimate public obligation. Everything else imposes privacy, interpretation, and maintenance costs.
Calibration is a human infrastructure requirement. If educators or supervisors apply criteria, they need shared examples, opportunities to compare judgments, and processes for resolving disagreement. Reliability should not mean eliminating judgment by reducing every performance to trivial elements. It means making judgment disciplined, inspectable, and improvable.
The learner should be able to see the claim, criteria, evidence, and interpretation. That transparency supports self-regulation and makes error correction possible. It also changes the tone of measurement. Evidence becomes something a learner can use and challenge, not a hidden extraction that returns only a verdict.
Matching evidence to risk
Low-consequence decisions can tolerate speed and uncertainty. A recommendation for the next practice item should be easy to override and cheap to correct. High-consequence decisions require stronger coverage, independent review, accessibility, security, documentation, and appeal. An institution reverses this logic when it collects exhaustive low-value behavioral data while relying on one compressed score for a life-changing gate.
A risk-proportionate design asks about both kinds of error. What happens if an unready person is approved? What happens if a ready person is excluded? Safety-critical settings rightly attend to the first, but gatekeeping systems often ignore the cumulative harm of the second. Multiple routes and supervised opportunities can protect the standard while reducing false exclusion.
The purpose of measurement is not to make a person fully knowable. It is to support a bounded claim well enough for a bounded decision. That modesty is the beginning of precision rather than its enemy.
Measuring Learning and Assessment Models
Limitations of Traditional Assessment Methods
An assessment can be useful for one decision and inadequate for another. A standardized test may sample a broad domain consistently while telling a teacher little about the reasoning behind a particular error. A classroom project may reveal that reasoning while providing too little comparable evidence for a system-level estimate. Neither becomes valid simply by being traditional or digital. The testing standards locate validity in the evidence for a proposed interpretation and use, not in an instrument’s brand. 5
Consider a report containing only a percentile. It describes standing in a reference distribution, not which concepts the learner can explain or how performance changed. My recommendation is to pair such a report with information relevant to the next instructional decision, while making clear that the additional information answers a different question.
Introduction to the Growth Assessment Model
A growth model represents change between observations under specified assumptions. It does not discover an objective quantity of merit inside a learner. The tasks, scale, time interval, supports, and missing observations determine what the estimated change means. Two learners may show the same score change after very different opportunities to practice.
A proposed platform might display a trajectory through a skill sequence. Before calling that trajectory growth, ask whether later tasks are comparable, whether hints were used, and whether the scoring rule changed. Advancement through software is a record of software performance. A delayed task outside the sequence can provide additional evidence about the capability the sequence was intended to develop.
Role of Data and Real-Time Feedback
Feedback has a stronger case than the claim that speed alone improves learning. Shute’s review describes formative feedback as information intended to modify thinking or behavior, with effectiveness depending on its form, timing, task, and learner. 6 An immediate answer can correct a misunderstanding, but it can also supply the very work an assessment was meant to observe.
For a hypothetical mathematics task, separate the first response, the hint, the revision, and a later unassisted problem. That record lets a teacher recognize improvement without misreporting assistance as independent mastery. The purpose of immediacy is to make an appropriate next action possible, not to make every judgment final sooner.
The Future of Academic Testing and Assessment
The Need for Change in Testing Methods
My case for assessment reform is a case for better alignment between evidence and consequence. Replacing an examination with a portfolio is not enough. Specify which capability the examination failed to represent, how the portfolio will elicit it, and how accessibility, authorship, scoring, and appeal will be handled.
A challenge assessment could offer another route to progression while retaining the same public standard. Its success should be judged by the quality of the decisions and by who can use the route, not by how unconventional it appears. A new gate can be just as arbitrary as the one it replaces.
Technology-Enhanced Assessments
An adaptive assessment selects subsequent tasks using earlier responses and a measurement model. That selection may improve the efficiency of gathering evidence, but validity still depends on the item pool, calibration, population, purpose, and consequences. A precise numerical estimate does not certify adequate coverage of the domain. 5
In a proposed adaptive algebra assessment, a teacher should be able to inspect what was sampled and what was not. If the estimate will determine placement, provide a way to demonstrate an unexpectedly strong capability, review accessibility problems, and request a second form of evidence. The system’s confidence is one input to the decision, not authority to close the route.
Expansion on Real-Time, Low-Stakes Testing
Low stakes describe consequences, not screen design. A quiz is not low stakes if every error enters a permanent risk score or influences placement without the learner’s knowledge. Nor does an interactive format guarantee freedom from anxiety.
Use brief practice checks when they serve an instructional purpose. Explain whether errors are retained, who sees them, and when a separate readiness judgment will occur. The design proposal is to make practice safe enough to reveal misunderstanding; that is a governance choice, not a claim that every frequent digital quiz reduces stress.
Importance of Regular, Low-Stakes Testing
Yao and colleagues’ 2024 meta-analysis combined 118 K–12 studies and reported a positive average achievement effect for broadly defined formative assessment. The reported overall effect was small; the synthesis does not identify one universally superior formative-assessment design. That supports using formative evidence thoughtfully, not promising that any schedule of quizzes will reproduce an average effect. 7
Frequency should follow the learning question. A brief check may be valuable after a new explanation; a delayed check answers a different question about retention. Neither replaces an opportunity to respond to feedback. Counting assessments without examining what changed afterward confuses monitoring with support.
The Future of Testing and Assessments
The future worth pursuing is an assessment system that can explain its judgments and correct them. That goal may use digital tools, trained observers, work samples, or all three. It does not require the maximum possible stream of learner data.
For a proposed pilot, predefine both errors of concern: advancing a learner whose evidence does not support readiness and holding back a learner who can demonstrate it. Review each error with the learner and the decision owner. A technology that makes one error less visible has not necessarily made assessment better.
Conclusion
Embracing a Merit-Based Framework in Learning
In this book, merit is a normative commitment to recognizing demonstrated capability under described conditions. Analytics does not validate that commitment by existing. Leaders must still decide which capability matters, whether opportunities to demonstrate it are fair, and what burden a learner must carry to challenge an error.
The evidence argument therefore has two tests. Does the measure support the interpretation? Is the resulting decision justified for the person affected? Treating those questions separately protects the book’s central distinction between evidence and permission.
The Benefits of Data Analytics in Education
Introduction to Data Analytics in Education
Learning analytics organizes records so people or systems can examine patterns. A pattern may help identify a question worth investigating: an unexpectedly common error, missing access to a resource, or a group that receives little feedback. It does not, by itself, identify the cause.
Two research syntheses illustrate the distinction. Liu and colleagues’ 2025 meta-analysis of 34 studies reported positive average outcomes from analytics-based interventions, with substantial heterogeneity and evidence of publication bias in the knowledge-acquisition subset. Kaliisa and colleagues’ dashboard review found weak causal evidence for achievement improvements and problems such as comparing self-selected users with nonusers. They examined different intervention scopes. Neither establishes that displaying data automatically improves learning. 89
Applications of Data Analytics in Education
Separate three applications: describing recorded activity, predicting an outcome, and evaluating an action. Login frequency can describe access to a platform. It is not a direct measure of attention. A prediction of noncompletion may be accurate while the recommended outreach is ineffective. An evaluation of outreach must observe what happens when that support is offered.
A hypothetical district could notice students with missing platform activity and ask whether devices, credentials, accessibility, or scheduling prevented participation. Contact should begin with a question, not a label of disengagement. The absence of a digital trace is also compatible with learning elsewhere.
Data Analytics and Personalized Learning
Personalization should name the thing changed: task difficulty, pacing, representation, feedback, or human support. An apparent difficulty in geometry does not establish weakness in every spatial task, and a strong algebra score does not prove readiness for every advanced topic.
A proposed recommendation can be tested through the learner’s explanation and a relevant task. Preserve an override and an upward route. The educational claim concerns whether the changed support helps the learner, not whether a model can sort the learner into a profile.
Monitoring Student Performance and Progress
A time series is interpretable only when its conditions are visible. Record changes in task difficulty, scoring, attendance opportunities, assistance, and assessment coverage. Improvement on easier questions cannot carry the same claim as improvement on comparable work.
In a hypothetical classroom, a rising weekly score could prompt the teacher to offer an unfamiliar problem and ask the learner to explain the strategy. The response supplies new evidence. It would be unwarranted to infer dedication, effort, or a stable personal trait from the score trend alone.
Informing Curriculum Development and Improvement
A curriculum review can use results to locate trouble, but comparisons across units are not automatically comparisons of teaching methods. Topic difficulty, prior knowledge, assessment alignment, time, and staff support may differ.
If a simulation unit has higher scores than a lecture unit, first examine those differences. A stronger evaluation would compare credible alternatives for the same learning goal and assess more than familiarity with the taught format. The resulting decision could retain the simulation, revise it, or prefer another approach; the dashboard should not decide the answer in advance.
Enhancing Teacher Performance and Professional Development
Student outcomes matter in examining teaching, but they do not isolate a teacher’s contribution without assumptions about prior attainment, assignment, attendance, other instruction, and measurement. My recommendation is to combine observations, work samples, contextual information, and a discussion of practice before deciding what support is appropriate.
A teacher whose class shows improvement can explain what changed and invite colleagues to examine the work. That is a useful professional inquiry. It is not proof that copying the teacher’s method will cause the same result elsewhere. Coaching evidence provides a better basis for designing sustained support than an unnamed success anecdote. 10
Ethical Considerations and Challenges
Analytics governance should define purpose, access, correction, retention, and prohibited reuse. Encryption protects one aspect of a record; it cannot make an invalid inference fair or a needless collection necessary. These are separate responsibilities. 11
A hypothetical school could publish a short explanation of what each indicator means, allow learners to correct errors, and record who is responsible for consequential action. This is an institutional design proposal. It is not a claim that a privacy notice, secure login, or annual transparency report guarantees trust.
Future Trends in Educational Data Analytics
More complex prediction is a technical possibility, not an inevitable educational improvement. Any proposed feature should identify a decision that becomes better and the evidence required to establish that improvement.
Before adopting a risk score, ask whether staff can act on it, whether the action is beneficial, and whether a simpler request for help would reach the same learners. A system that finds more risk than the institution can responsibly address may add records without adding support.
Data’s Role in Decision Making and Learning Enhancement
Data-Driven Decision Making in Education
Data-informed decision making uses observations to examine instructional and strategic choices. It does not remove judgment from the selection of measures, comparison groups, or consequences. State the decision and the assumptions before inspecting a favorable chart.
If graduation rates rise after a program begins, consider changes in admissions, student composition, grading, local conditions, and concurrent support. The improvement is worth investigating, but timing alone does not establish causation. Report what is observed separately from what is inferred and what the institution proposes to do next.
Supporting Personalized Learning Through Data
Data can guide a question about the next useful task. It cannot assign a validated learning style or guarantee that a recommendation is appropriately challenging. Learner preferences are relevant to access and usability without becoming a prescription for one permanent teaching format.
In a proposed lesson, use an initial response to offer additional explanation or a more demanding task, then inspect the next response. Keep the change reversible. Personalization is responsible when new evidence can correct it, including evidence that the original recommendation underestimated the learner.
Identifying and Addressing Learning Gaps
A learning gap is a difference between the capability sought and the evidence currently available. An incorrect response does not distinguish missing knowledge from inaccessible directions, unfamiliar language, or a misleading question. Formative assessment is useful when investigation changes teaching or learner action. 2
For a hypothetical quadratic-equation lesson, ask learners to explain their chosen method before assigning extra practice. After support, use a different problem and a delayed check. The design creates a way to investigate improvement without inventing a successful outcome.
Improving Student Outcomes
An outcome should be specified before a program is judged: retention of knowledge, independent performance, access, attendance, completion, or another legitimate purpose. These outcomes need not move together. Improvement in one should not be reported as evidence for all.
If reading scores rise after a curriculum change, examine baseline differences, attendance, assessment changes, concurrent teaching changes, and an appropriate comparison. Publish an unfavorable or uncertain result with the same care as a favorable one. The institution’s commitment should be to discovering whether the intervention helps, not to protecting the purchase.
Utilizing Data for Resource Allocation and Performance Evaluation
Data-Driven Resource Allocation
Allocation combines evidence with values. A low score may indicate a need for support, but deciding which support deserves funding requires evidence about the available responses, their costs, and the people they reach. No analytics system determines the right allocation from a ranking alone.
A hypothetical district might compare tutoring, additional teacher planning time, accessible materials, and software for the same identified problem. Count implementation and ongoing support costs. Record which alternatives were rejected and why so the decision can be revisited rather than inherited as an unexplained budget line.
Evaluating Teacher Performance
Teacher evaluation should not convert a model estimate into a complete personal judgment. A result is produced within assignments, resources, curriculum, attendance, and measurement conditions. Where evidence is consequential, its interpretation requires scrutiny and a route to challenge errors.
Separate a professional-development conversation from a disciplinary decision. The first can explore tentative signals and alternative explanations; the second carries a greater burden of evidence and procedure. A school should not advertise a developmental dashboard and quietly reuse its uncertain indicators as a final employment verdict.
Predictive Analytics in Education
The Early Warning Intervention and Monitoring System provides a real example of an evaluated support process. In a randomized study of 73 high schools in three U.S. Midwest states, the first-year program reduced chronic absence and course failure. It did not establish improved on-time graduation, and several other measured outcomes showed no detectable benefit. The intervention combined identification, support assignment, and monitoring; it was not a test of a predictive algorithm alone. 12
The leadership implication is to evaluate the whole response. A prediction can initiate a conversation, but the quality, availability, and consequences of the support determine whether the institution has helped. Retain uncertainty in the learner’s record and allow direct evidence to challenge the predicted trajectory.
Assisted practice and independent proof
Feedback should improve the next attempt, not blur what the learner can do alone. Formative-feedback research supports information that is timely, specific, and usable, but the amount and timing of help alter what a successful response proves. 6 A worked hint, pronunciation model, calculator, peer explanation, or AI suggestion may be entirely appropriate during learning. Evidence records should make that support visible.
A useful mastery sequence distinguishes at least four conditions: performance with ordinary learning supports, performance after feedback, delayed performance, and performance under the degree of independence relevant to the decision. “No help” is not a virtue in every lesson, and support is not cheating by definition. The problem arises when an assisted practice score is reported as independent mastery.
Automated oral-reading assessment illustrates the validation burden. Researchers have tested speech-enabled fluency measures and automatic oral-reading accuracy against human or diagnostic criteria in bounded settings. 1314 That is evidence that a particular measurement approach can be evaluated, not permission to treat any speech model as a diagnostician.
Before use, leaders must name the language, age range, reading task, microphone conditions, reference raters, acceptable error, and consequence of a mistake. Accent, dialect, speech difference, disability, background noise, and text difficulty can change performance. Automated scoring may help a trained adult review more samples or notice a pattern. It should not independently determine disability status, placement, retention, access, or discipline.
Foundational reading guides also clarify that a score must remain attached to a component. Phonemic awareness, decoding, fluency, vocabulary, and comprehension interact but are not interchangeable. 1516 A system that reports one “reading level” without showing which evidence produced it has compressed away information needed for instruction.
Human judgment is also an instrument
Rejecting automated certainty does not make human scoring automatically valid. Raters can disagree, drift over time, import bias, overlook context, or reward surface features unrelated to the construct. The response is not to remove judgment from every serious performance. It is to design judgment so its grounds can be examined.
A defensible scoring process names the construct, supplies examples at different levels, trains raters on disagreements, samples double-scored work, and checks that the same relevant criteria are applied consistently across raters, tasks, and formats. When the work is complex, the record should preserve comments or criterion-level decisions rather than only a total. A learner can then understand what evidence led to the judgment and what improvement would look like.
Human and automated scoring can be compared, but agreement alone is not enough. Two systems can agree because both reward length, conventional language, or familiar solutions. Validation asks whether the score supports the intended interpretation and decision. Review disagreements, unusual responses, accessibility cases, and the consequences of error. If a model is used to prioritize human attention, test whether the prioritization causes some learners’ work to receive less careful review.
Delayed and changed conditions reveal different things
Immediate post-tests are often useful for detecting whether instruction produced an initial change. They are weak evidence of durability by themselves. A delayed check shows what remains after some forgetting and intervening experience. A changed task examines whether the learner can identify the relevant structure rather than reproduce a practiced surface.
Independence is also graduated. A professional may appropriately use references, calculators, code tools, colleagues, or AI. Evidence should reproduce the conditions in which competent performance matters. At the same time, some component capabilities must remain available when a tool is wrong, unavailable, or unsafe. The assessment program should identify those components explicitly rather than declaring every authentic tool either permitted or forbidden.
This creates an evidence portfolio with interpretable layers: supported practice, revised performance after feedback, delayed retention, transfer to changed conditions, independent components, and authentic tool-enabled production. No layer is universally superior. Together they prevent one convenient score from claiming more than it observed.
When two measures disagree
Disagreement is not an inconvenience to average away before anyone sees it. Consider a hypothetical learner who performs well in a project defense but poorly on a timed component test. The institution should first ask whether the two tasks support the same claim. One may require explanation with references, another quick retrieval without them. If the intended responsibility includes both, the disagreement describes a profile that deserves attention. If speed is irrelevant to the responsibility, the timed score may introduce a burden the institution has not justified.
The reverse disagreement is also informative. A learner may execute familiar items quickly but struggle to decide which method applies in a changed problem. Calling the first score mastery of the whole domain would conceal the distinction. Calling it worthless would discard valid evidence of a narrower capability. The task of an evidence system is to preserve what has been shown while identifying the next question.
A third measure is not automatically the solution. Three tests with the same narrow format may repeat the same omission. A richer body of evidence should vary the conditions that matter to the claim: explanation, independent execution, unfamiliar context, delayed performance, or authentic tool use. Variation should be purposeful. Collecting a large portfolio without deciding what each artifact contributes merely moves the interpretation problem into a larger archive.
Now suppose two raters disagree about the same work. Begin with their reasons and criteria. They may have noticed different evidence, interpreted an ambiguous criterion differently, or made an error. The response could involve clarifying the rubric, comparing examples, reviewing the work together, or obtaining another qualified judgment. It should not automatically punish the learner for disagreement produced by the assessment process.
The decision owner must also say how multiple observations will be combined. An average can conceal failure on an essential safety component; a strict minimum on every component can exaggerate the consequence of one noisy observation. Neither aggregation rule is a natural fact. State which capabilities are essential, which can compensate for one another, and which require more evidence before a decision. Keep that rule visible before the learner attempts the task.
The same caution applies to change over time. Moving upward in a rank does not necessarily mean learning increased: the reference group may have changed. Learning can increase without a rank changing if others improved too. Likewise, moving from one rubric category to another does not establish that each category represents an equal amount of growth. A numerical code attached to a category should not silently acquire the properties of a measurement scale.
For a hypothetical progress report, distinguish the observed work, the category assigned, the comparison used, and the claim about change. If tasks became more difficult, explain why scores can be compared. If assistance changed, report it. If the assessment was replaced, acknowledge a break in the series unless there is a defensible linking method. These are requirements for an interpretable account, not an instruction to avoid measuring growth.
Finally, decide what the learner can do with the disagreement. A useful report identifies what remains uncertain and what additional opportunity could resolve it. It does not demand endless proof without explaining when enough evidence will be enough. A bounded decision can acknowledge uncertainty while granting an appropriately bounded responsibility. That is more precise than either unconditional confidence or permanent refusal.
This worked example is the author’s proposed method for reviewing evidence. It reports no trial outcome and promises no automatic improvement in fairness. Its purpose is to make the institution’s reasoning inspectable: the capability, the task, the observation, the interpretation, the rule for combining evidence, and the consequence should remain distinguishable throughout the decision.
A minimum viable evidence model
For each important capability, define the claim, acceptable evidence, conditions, criteria, feedback cycle, and decision owner. Preserve representative artifacts rather than only scores. Record enough context to interpret the observation, including supports and degree of independence. Ask for repeated evidence when the consequence is high or the performance is variable.
The model should be economical. More data can make a system less interpretable and more invasive. Collect the evidence needed to improve or decide, retain it only as long as its purpose requires, and return it in a form the learner can understand.
Implications for leaders
- Report growth, mastery, and transfer separately.
- Define domains and criteria before selecting instruments.
- Ensure formative evidence produces a specific action pathway.
- Preserve access to representative work behind aggregate indicators.
- Match the amount and quality of evidence to the consequence of the decision.
Questions to carry forward
- Which claim does each current measure actually support?
- Can learners see why a judgment was made and what evidence would change it?
- How far beyond the assessment context is your institution claiming transfer?
Notes
- 1Barnett, Susan M. and Ceci, Stephen J.. 2002. “When and Where Do We Apply What We Learn? A Taxonomy for Far Transfer”. Psychological Bulletin, vol. 128, no. 4, 612–637.↩
- 2Black, Paul and Wiliam, Dylan. 1998. “Assessment and Classroom Learning”. Assessment in Education: Principles, Policy & Practice, vol. 5, no. 1, 7–74.↩
- 3Hattie, John and Timperley, Helen. 2007. “The Power of Feedback”. Review of Educational Research, vol. 77, no. 1, 81–112.↩
- 4National Research Council. 2011. “Incentives and Test-Based Accountability in Education”. The National Academies Press.↩
- 5American Educational Research Association, American Psychological Association, National Council on Measurement in Education. 2014. “Standards for Educational and Psychological Testing”. American Educational Research Association.↩
- 6Shute, Valerie J.. 2008. “Focus on Formative Feedback”. Review of Educational Research, vol. 78, no. 1, 153–189.↩
- 7Yao, Yuankun, Amos, Michelle, Snider, Karrie, et al.. 2024. “The impact of formative assessment on K-12 learning: a meta-analysis”. Educational Research and Evaluation, vol. 29, no. 7-8, 452-475.↩
- 8Liu, Yan, Wang, Wei, Xu, Enwei. 2025. “The Effectiveness of Learning Analytics-Based Interventions in Enhancing Students’ Learning Effect: A Meta-Analysis of Empirical Studies”. Sage Open, vol. 15, no. 2.↩
- 9Kaliisa, Rogers, Misiejuk, Kamila, López-Pernas, Sonsoles, et al.. 2024. “Have Learning Analytics Dashboards Lived Up to the Hype? A Systematic Review of Impact on Students' Achievement, Motivation, Participation and Attitude”. Proceedings of the 14th Learning Analytics and Knowledge Conference, 295-304.↩
- 10Kraft, Matthew A., Blazar, David, Hogan, Dylan. 2018. “The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence”. Review of Educational Research, vol. 88, no. 4, 547–588.↩
- 11National Institute of Standards and Technology. 2020. “NIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0”. NIST.↩
- 12Faria, Ann-Marie, Sorensen, Nicholas, Heppen, Jessica, et al.. 2017. “Getting students on track for graduation: Impacts of the Early Warning Intervention and Monitoring System after one year”. Institute of Education Sciences, U.S. Department of Education, REL report, no. 2017-272.↩
- 13van der Velde, Max, Harmsen, Wieke, Veldkamp, Bernard P., et al.. 2025. “Speech Enabled Reading Fluency Assessment: a Validation Study”. International Journal of Artificial Intelligence in Education, vol. 35, 2569-2595.↩
- 14Molenaar, Bo, Tejedor-Garcia, Cristian, Cucchiarini, Catia, et al.. 2023. “Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics”. Interspeech 2023, 5232-5236.↩
- 15National Reading Panel. 2000. “Teaching Children to Read: An Evidence-Based Assessment of the Scientific Research Literature on Reading and Its Implications for Reading Instruction”. National Institute of Child Health and Human Development.↩
- 16Foorman, Barbara, Beyler, Nicholas, Borradaile, Kelley, et al.. 2016. “Foundational Skills to Support Reading for Understanding in Kindergarten Through 3rd Grade”. National Center for Education Evaluation and Regional Assistance, Institute of Education Sciences, U.S. Department of Education, no. NCEE 2016-4008.↩