Learning Without Gatekeepers LWG Start reading

Part III · Chapter 10Data With Dignity

0%

Part III · Chapter 10

Data With Dignity

The ability to collect a signal is not permission to keep it, infer from it, or use it against a learner.

18 minute read 4,071 words Revised August 2026

Digital learning systems can record detailed interaction traces and, in some research settings, combine them with gaze or physiological signals. Reviews describe these technical possibilities without establishing that every school collects every signal or that collection measures learning. 12 The technical ability to observe creates a temptation to treat visibility as knowledge. But a signal becomes educational evidence only through interpretation, and interpretation becomes legitimate only through purpose, validity, proportionality, and governance.

Data with dignity begins by reversing the default question. Do not ask, “What can the platform collect?” Ask, “What decision or learning action requires evidence?” Then collect the least intrusive information capable of supporting it. The aim is to reduce unnecessary exposure and force a clearer account of the decision. When teams cannot name the action a field will change, the field is probably administrative residue or speculative surveillance.

The NIST Privacy Framework treats privacy as an organizational risk-management problem involving data processing, governance, communication, and control, not merely a notice displayed to users. 3 For learning systems, that means mapping who collects information, who can infer from it, how long it persists, what other datasets can be joined, and what consequence can follow. Consent is weak when participation in school or employment is not meaningfully optional.

Biometric and behavioral data raise special concerns. Eye tracking can support accessibility research and investigation of visual attention under controlled conditions. It cannot directly reveal a learner’s thoughts, motivation, honesty, or understanding. Gaze depends on task, display, vision, fatigue, and many other factors. Turning it into a continuous “focus score” introduces inference on top of inference while normalizing observation of a learner’s body. The stronger the claim and consequence, the less adequate a correlational signal becomes.

Prediction raises a relational design risk even when it is accurate on average. If a system predicts that a learner is likely to fail, adults might offer support or lower expectations, reduce challenge, and communicate doubt. A feedback effect in which the response helps produce the forecast outcome is a possibility to investigate, not an established result for every system. Leaders must distinguish decision support from decision replacement and test whether interventions triggered by a prediction actually help. Accuracy without a beneficial action is not educational value.

Data also move. A record collected to personalize practice can later be used for discipline, placement, admissions, marketing, or model training. Purpose limitation must therefore be technical and contractual, not aspirational. Access logs, retention schedules, deletion, portability, vendor restrictions, and human review are features of instructional integrity. Learners should know what is collected and be able to see and correct the records that affect them.

Dignity also requires the right not to become fully legible. Education should cultivate private reflection, experimentation, and identity development. Not every draft belongs in a permanent profile. Not every interaction needs to improve a model. A system that remembers everything can make risk-taking irrational. Protected forgetting is part of a healthy learning environment.

For leaders, the standard is not data minimization at any cost. Some information is necessary to check equality of treatment, diagnose learning difficulties, provide task-relevant support, and protect safety. The standard is a documented relationship among purpose, evidence quality, intrusion, access, retention, and consequence. Data should serve the learner’s development and legitimate public responsibility, not the institution’s appetite for control.

Privacy is a limit on institutional appetite

Education organizations have historically kept records because instruction and public administration require memory. Digital systems change the scale, granularity, persistence, and combinability of that record. A workbook response can become a timestamped event joined with device information, location, browsing behavior, message history, gaze estimates, and predictions. The fact that a system can combine these signals does not make the resulting profile educationally necessary or valid.

Privacy is sometimes framed as a preference to be balanced against innovation. For learners, it is also a condition of agency. A learner may reasonably treat exploration differently when uncertainty, search, pauses, errors, or emotional expression could become a durable judgment. That is a design risk to investigate, not a universal behavioral law. Protected exploration is part of learning. A system that cannot forget can make an early struggle follow a learner long after the underlying capability has changed.

The NIST Privacy Framework organizes privacy risk as a governance and engineering responsibility rather than a notice form. It emphasizes identifying data processing, governing it, controlling it, communicating about it, and protecting it. 3 For education leaders, that begins before procurement. What decision requires the data? Why is a less intrusive signal insufficient? Who can use it? How long does it survive? What inference is prohibited? What remedy exists when the record or inference is wrong?

Data minimization is an educational design tool

Minimization is not simply deleting fields after collection. It means designing the learning and evidence process so unnecessary data is never created or centralized. If a learner needs feedback on practice, the system may not need a permanent event-level history. If an institution needs evidence of completion, it may not need every keystroke used to produce the work. If a teacher needs to identify a misconception, a local summary may be sufficient without exporting identifiable behavior to a vendor.

A disciplined data inventory attaches each field to a purpose, decision owner, retention period, access group, and deletion trigger. “May be useful later” is not a purpose. Product improvement and educational evaluation are separate purposes that require separate justification. Training a commercial model is another purpose again. Consent to learn in an institution should not be treated as consent to every downstream use.

Minimization can help make purpose explicit. An analyst can find correlations in a large event stream without first stating an educational claim. The resulting model may predict an institutional outcome while offering no valid explanation of learning. My recommendation is to prefer evidence aligned with a defined decision over additional fields collected without that justification.

The special danger of inferred states

Some systems attempt to infer attention, emotion, identity, honesty, risk, or ability from faces, voices, bodies, interaction patterns, or physiological signals. These claims deserve heightened scrutiny. Internal states are not directly observable, expression varies across people and cultures, and context can dominate the signal. Even an accurate sensor does not guarantee an accurate psychological inference, and an average association does not justify a consequential judgment about one learner.

Gaze is an example. Eye movement can be useful in controlled research and usability analysis. It does not provide a universal meter of attention or understanding. A learner may look away while thinking, reread because a passage is profound rather than confusing, or hold steady gaze while mentally disengaged. Disability, neurodivergence, fatigue, culture, equipment, lighting, and interface design can change the measurement. Turning such a signal into a compliance score layers interpretive uncertainty onto a power imbalance.

Emotion recognition raises related concerns. Barrett and colleagues’ review challenges reliable, context-independent inference of specific emotions from facial movements. It does not imply that expressions contain no information, and it is not a review of every vocal classifier. A face is not a transparent display of an inner category. 4 When a school treats a probabilistic inference as fact, the learner may have no practical way to disprove it. The safest default is not to use biometric or affect inference for consequential educational decisions. If a narrowly bounded use is proposed, the burden belongs to the institution to demonstrate necessity, validity in the actual population and context, proportionality, accessibility, security, and a meaningful non-biometric alternative.

Surveillance can invalidate its own evidence

Observation may change the conditions under which evidence is produced. As a design risk, a learner could optimize for visible compliance, avoid an ambiguous action, or seek a way around a monitor. A system that interpreted such adaptation as a stable trait would be making an additional, unvalidated inference. The point is a threat to examine, not a claim that every learner responds to surveillance in the same way.

Remote proctoring illustrates the tension. Institutions may have a legitimate need to protect assessment integrity. But room scans, continuous video, automated behavior flags, and invasive device controls can create unequal burdens, expose family and living conditions, and produce accusations that are difficult to contest. Leaders should first redesign the assessment: use open-resource tasks, oral follow-up, varied prompts, staged work, authorship checks, and sampling over time. Surveillance should not compensate for an assessment whose validity depends on pretending authentic tools do not exist.

Governance for educational data

A credible governance structure assigns named responsibility. Academic or instructional owners define the learning claim. Assessment expertise evaluates validity. Privacy and security expertise examine collection and threat. Accessibility and civil-rights expertise examine burden and differential impact. Learners and educators describe how the system changes behavior in practice. Procurement translates the limits into enforceable vendor terms. A review body can stop or revise use when evidence fails.

The NIST AI Risk Management Framework similarly treats governance as continuous and connects it with mapping context, measuring risk, and managing it over time. 5 UNESCO’s guidance on generative AI in education emphasizes a human-centered approach, data privacy, age appropriateness, and validation of pedagogical use. 6 These frameworks do not decide local policy, but they make one point clear: launching a tool is the beginning of oversight, not the end.

Vendor contracts should specify data purpose, location, subprocessors, access, retention, deletion, incident reporting, model training, secondary use, audit rights, and what happens when the relationship ends. Institutions need an export and deletion path that does not depend on vendor goodwill. A free or low-cost product can impose a high future cost if data and workflow cannot be recovered.

The learner’s rights in an evidence system

The following are governance commitments proposed by this book, not a claim that every jurisdiction already guarantees each right. Applicable law and institutional obligations require separate review.

A dignity-preserving system gives the learner practical capabilities, not merely a long notice. The learner can see what evidence exists, understand the claim it supports, correct factual errors, add context, challenge an inference, know who accessed it, and obtain relevant evidence in a usable form. They can also practice without every attempt becoming consequential.

Deletion rights must be reconciled with legitimate records, public obligations, research integrity, and safety. The answer will vary by context. The institution should state the exception narrowly rather than using a possible future obligation to justify indefinite retention of everything.

Special care is required when learners cannot freely refuse. In compulsory education, employment-linked training, or access to essential services, a consent checkbox does not erase the power imbalance. Leaders should provide a functionally equivalent, non-punitive alternative when optional high-risk data collection is proposed. If no equivalent alternative is possible, the collection should be treated as mandatory and justified under the stricter standard that implies.

A restrained technical architecture

Architecture can embody these principles. Keep private reading or practice state on the device when synchronization is unnecessary. Separate identifying data from learning evidence. Aggregate only after determining that aggregation preserves the needed meaning. Encrypt data in transit and at rest. Restrict access by role and purpose. Log consequential access and changes. Set deletion automatically. Test restoration and exit before launch.

Where data must travel, use selective disclosure. A learner should be able to present evidence for a capability without exposing an entire educational history. Provenance can show who observed or verified an artifact, when, under what conditions, and against which criteria. It should not become a universal identity graph available to every institution.

Local-first design has limits. A device can be lost, storage can be cleared, and learners may need access across devices. That is why export and validated import matter. The broader principle is not that every record must remain local. It is that centralization requires a purpose strong enough to justify the additional power and risk.

Data with dignity does not reject evidence. It insists that the person remains larger than the record. The institution may make a bounded decision from bounded evidence. It does not acquire a general right to observe, predict, or preserve the learner indefinitely.

Enhancing Parental Involvement and Accountability Through Data

Strengthening Parental Involvement

A parent portal is a communication channel, not an intervention with an automatic achievement effect. It may make missing work visible, but whether a family can respond depends on language, time, access, relationships, and the kind of help requested. Research on parental involvement distinguishes forms of involvement and often relies on associations rather than causal comparisons. 78

My recommendation is to make messages actionable and proportionate: what was observed, what remains uncertain, what support the school will provide, and how the family can add context. Do not turn a child’s every hesitation into an alert. A family should not have to reproduce a surveillance dashboard at home to be considered supportive.

Benchmarking and Standardization

Benchmarking compares results with a reference; standardization makes specified procedures more consistent. Neither establishes fairness by itself. A reference group may be inappropriate, a task may include irrelevant barriers, or a consistent procedure may support the wrong decision.

Before sending a comparison to families, explain the population, date, purpose, and uncertainty. A percentile is not a percentage of the curriculum mastered. Where the question is an individual learner’s next step, provide direct work and criteria rather than allowing a rank to stand in for an instructional explanation.

Promoting Accountability and Transparency

Transparency should expose how a decision was made, not merely publish favorable totals. State the measure, exclusions, missing data, changed definitions, and reasonable alternative interpretations. Describe what the institution will do in response and who can question it.

A hypothetical district report could show both opportunities offered and outcomes observed, explain where the two cannot yet be causally connected, and include the response to an unsuccessful pilot. The recommendation is to make institutional judgment inspectable. Publishing more information does not itself prove greater trust, better teaching, or improved learning.

The Importance of Data in Specialized Training Programs

Learning Measurement in Specialized Training

Where a trainee’s future work carries safety consequences, define the capability and the limits of the evidence especially carefully. Performance in a simulation establishes what happened in that simulation. It does not automatically establish readiness in a clinical encounter, emergency, or legally consequential situation.

Consider a proposed professional-training exercise involving a difficult decision under pressure. Record the scenario, available information, assistance, scoring criteria, and rater judgment. A debrief can examine reasoning that a timing metric missed. This is a design example, not a report that an unnamed academy improved performance. Real-world transfer requires evidence in conditions relevant to the responsibility being granted. 9

Innovative Approaches to Learning Measurement

A simulation can make selected actions observable and repeatable. A log may record tool movement, sequence, timing, or a response to a scripted event. Those records should not be renamed general competence without a validity argument.

For a hypothetical surgical-skills simulator, distinguish instrument accuracy from clinically meaningful assessment. A trainee could become faster by omitting a necessary check. An automated score might miss judgment about when not to proceed. Qualified supervisors should determine which dimensions the simulation can represent and which require other evidence. No clinical efficacy claim for a particular simulator is made here.

The Future of Data-Driven Training

A defensible training system should accumulate evidence relevant to the next responsibility while keeping its uncertainty visible. Adaptation can offer more practice after an error, but certification should depend on the required performance standard rather than completion of a personalized sequence.

In a proposed cybersecurity exercise, separate success with hints from response to an unfamiliar incident. Record how assistance changed the result and ask the learner to explain the decision. The institutional question is whether the evidence supports the intended responsibility. Neither an AI label nor a detailed event log answers that question alone.

Conclusion

Embracing Data-Driven Education

The case for data is conditional: collect it when it supports a warranted decision or useful feedback, and stop when it adds intrusion without a defensible purpose. NIST’s privacy framework informs this governance approach; it is not research proving that analytics raises achievement. 3

An institution should be prepared to discontinue a dashboard that offers no actionable information, even if its reports look sophisticated. Conversely, a modest record of work and feedback may be valuable when it helps a learner understand and challenge a judgment. The standard is the quality of the use, not the quantity of the collection.

What a signal can and cannot say

Wearable and multimodal systems can collect streams of heart activity, motion, gaze, audio, video, skin response, and interaction events. Reviews document rapid technical expansion and a growing range of analytic methods. 21 This establishes feasibility. It does not establish that a signal measures the educational construct named in a dashboard.

Every inference has at least three layers. A sensor estimates a physical event. A model maps events to a construct such as stress, cognitive load, or engagement. An institution then maps that construct to an action. Error and disagreement can enter at every layer. Accuracy in detecting a heartbeat does not prove accuracy in classifying attention, and an association with a group average does not justify a decision about one learner.

HRV is not an attention meter

Heart-rate variability is often presented as an objective window into stress or cognitive demand. A systematic review of HRV in educational research found moderate support for some stress-related interpretations but uncertain attention validity and limited evidence connecting stress measures to performance. 10 A study in clinical reasoning found correlations among heart measures, some performance outcomes, and self-reported cognitive load in its particular task. 11 The second result does not cancel the first limitation. It shows why context-specific association must not become universal classification.

Heart measures vary with posture, movement, respiration, medication, illness, fitness, device fit, time of day, and emotional and physical conditions unrelated to instruction. 10 A school must not label a learner distracted, unready, dishonest, or at risk because an HRV model produced a score. If physiological data are used in voluntary research, the construct, uncertainty, exclusion criteria, retention, and non-participation route belong in the design before collection begins.

Blinks reverse the easy story

Blink research is useful precisely because it resists a universal rule. Blink patterns varied with engagement in scene content in one laboratory study. 12 Blink rate increased with cognitive load in an auditory oddball task that did not require visual attention. 13 In a later sentence-listening study, blinking decreased under more challenging auditory conditions. 14

These findings need not contradict one another. Blinking is shaped by visual demands, timing, task structure, fatigue, environment, and individual physiology. They do contradict the claim that more or fewer blinks transparently reveal engagement or load. A real-time “focus score” that ignores task-specific direction and uncertainty manufactures precision.

Gaze requires the same restraint. Eye tracking can describe where a person looked under measured conditions and can support aggregate usability research. It cannot tell why the learner looked there, what was understood, or whether looking away represented distraction, recall, planning, sensory regulation, or reflection. Gaze should not identify knowledge gaps or trigger consequential interventions without direct performance evidence.

Human-centred means governed by the people affected

A systematic review of human-centred learning analytics and educational AI found limited end-user involvement and recurring gaps involving human control, safety, reliability, and trust. 15 The result is a warning against treating a participatory workshop or friendly interface as proof that a system is human-centred. Learners and educators need authority before launch, access to understandable evidence during use, and a real ability to challenge or stop harmful operation.

Ethical principles for learning analytics emphasize transparency, learner control, security, accountability, and the institution’s duty to use data for legitimate educational purposes. 16 Those principles become meaningful only as operational rules: named decision owners, prohibited uses, access logs, deletion schedules, validation thresholds, incident processes, and equivalent non-biometric routes.

The default for biometric or behavioral mental-state inference in consequential education should remain non-use. A proposed exception must show that the decision is necessary, that direct and less invasive evidence is insufficient, that the measure is valid in the actual population and setting, and that a learner can refuse or contest it without penalty.

A validation protocol for intimate data

If an institution believes an intimate signal may serve a legitimate low-risk purpose, it should begin with a written claim rather than a device. The claim names the observable signal, the construct someone proposes to infer, the action that could follow, and the educational benefit expected from that action. Each link requires separate evidence. A pulse sensor may detect intervals accurately while the stress model performs unevenly; a stress estimate may correlate with a group measure while the proposed intervention provides no benefit.

The next step is a less-invasive-alternative test. Direct performance, learner request, teacher observation, a short voluntary check-in, or a change to the task may answer the educational question with less uncertainty and less power over the learner. Convenience for the institution is not enough to justify biometric collection. If a direct measure is available, the indirect intimate signal should ordinarily lose.

Validation must occur in the intended population and conditions. Device fit, skin contact, lighting, background noise, mobility, language, disability, medication, and classroom routine can change measurement. Overall accuracy is insufficient when the proposed decision concerns a particular learner under particular conditions. Report false positives and false negatives across relevant tasks, devices, and operating conditions without collecting personal attributes unrelated to the validation question. The analysis itself must satisfy necessity and privacy requirements.

The institution should then run a silent or advisory pilot. In a silent pilot, the output does not affect learners; researchers compare it with appropriate reference evidence and look for failure. In an advisory pilot, a trained person sees the output but retains responsibility, records whether it was useful, and documents disagreements. Automatic intervention should not be the first real-world test of a model whose errors are still being discovered.

Before any operational use, define a stopping rule. Excessive disagreement, differential error, learner distress, security failure, function creep, or absence of educational benefit should pause or end collection. A model should also expire. Changes to sensors, software, population, task, or institutional purpose trigger revalidation rather than inheriting trust from the previous version.

Finally, communicate uncertainty in the interface. A color, gauge, or rank can make a probabilistic estimate look like a fact. Show what was measured, what was inferred, the confidence and known limitations, the recency of the evidence, and the person responsible for the decision. Give learners a way to add context before action is taken. The purpose of a validation protocol is not to make surveillance respectable. It is to make weak or unnecessary uses fail before they acquire institutional momentum.

The dignity test

Before approving a data practice, ask five questions: Is the purpose specific? Is the signal valid for that purpose? Is there a less intrusive alternative? Can the learner understand and contest the use? What happens if the data are wrong, exposed, or repurposed? A practice that cannot survive those questions is not ready for deployment.

Governance should include learners and educators, not only technology, legal, and procurement staff. The people subject to collection often see contextual harms that a risk register misses. Their participation is not a substitute for institutional responsibility; it is part of competent design.

Implications for leaders

  • Require a named action and validity argument for every sensitive data field.
  • Ban high-consequence inferences from gaze, emotion, or other weak biometric proxies.
  • Separate practice records from permanent evaluative records.
  • Build access, correction, retention, deletion, and vendor-use limits into contracts and systems.
  • Evaluate whether predictive interventions improve outcomes rather than only whether predictions are accurate.

Questions to carry forward

  1. What does your institution collect because it can rather than because it should?
  2. Which records created for support can later become evidence against a learner?
  3. Where does a learner have protected space to be unfinished and unobserved?

Notes

  1. 1Mohammadi, Mehrnoush, Tajik, Elham, Martinez-Maldonado, Roberto, et al.. 2025. “Artificial Intelligence in Multimodal Learning Analytics: A Systematic Literature Review”. Computers and Education: Artificial Intelligence, vol. 8, 100426.
  2. 2Hong, Huaqing, Dai, Ling, Zheng, Xiulin. 2025. “Advances in Wearable Sensors for Learning Analytics: Trends, Challenges, and Prospects”. Sensors, vol. 25, no. 9, 2714.
  3. 3National Institute of Standards and Technology. 2020. “NIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0”. NIST.
  4. 4Barrett, Lisa Feldman, Adolphs, Ralph, Marsella, Stacy, et al.. 2019. “Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements”. Psychological Science in the Public Interest, vol. 20, no. 1, 1-68.
  5. 5Tabassi, Elham. 2023. “Artificial Intelligence Risk Management Framework (AI RMF 1.0)”. National Institute of Standards and Technology.
  6. 6Miao, Fengchun and Holmes, Wayne. 2023. “Guidance for Generative AI in Education and Research”. UNESCO.
  7. 7Wang, Xueshen and Wei, Yun. 2024. “The influence of parental involvement on students’ math performance: a meta-analysis”. Frontiers in Psychology, vol. 15, 1463359.
  8. 8Xu, Jianzhong, Guo, Shengli, Feng, Yuxiang, et al.. 2024. “Parental Homework Involvement and Students’ Achievement: A Three-Level Meta-Analysis”. Psicothema, vol. 36, no. 1, 1-14.
  9. 9Barnett, Susan M. and Ceci, Stephen J.. 2002. “When and Where Do We Apply What We Learn? A Taxonomy for Far Transfer”. Psychological Bulletin, vol. 128, no. 4, 612–637.
  10. 10Kim, Hyun Jin, Park, Yuyi, Lee, Jihyun. 2024. “The Validity of Heart Rate Variability (HRV) in Educational Research and a Synthesis of Recommendations”. Educational Psychology Review, vol. 36, 42.
  11. 11Solhjoo, Soroosh, Haigney, Mark C., McBee, Elexis, et al.. 2019. “Heart Rate and Heart Rate Variability Correlate with Clinical Reasoning Performance and Self-Reported Measures of Cognitive Load”. Scientific Reports, vol. 9, 14668.
  12. 12Ranti, Carolyn, Jones, Warren, Klin, Ami, et al.. 2020. “Blink Rate Patterns Provide a Reliable Measure of Individual Engagement with Scene Content”. Scientific Reports, vol. 10, 8267.
  13. 13Magliacano, A., Fiorenza, S., Estraneo, A., et al.. 2020. “Eye Blink Rate Increases as a Function of Cognitive Load During an Auditory Oddball Paradigm”. Neuroscience Letters, vol. 736, 135293.
  14. 14Coupal, Penelope, Zhang, Yue, Deroche, Mickael. 2025. “Reduced Eye Blinking During Sentence Listening Reflects Increased Cognitive Load in Challenging Auditory Conditions”. Trends in Hearing, vol. 29, 23312165251371118.
  15. 15Alfredo, R., Echeverria, V., Jin, Y., et al.. 2024. “Human-Centred Learning Analytics and AI in Education: A Systematic Literature Review”. Computers and Education: Artificial Intelligence, vol. 6, 100215.
  16. 16Pardo, Abelardo and Siemens, George. 2014. “Ethical and Privacy Principles for Learning Analytics”. British Journal of Educational Technology, vol. 45, no. 3, 438–450.

Private note

Add to your notebook

Notebook

Full-book search

Find an argument, source, or idea

Type at least two characters.