Responsible AI in Education Starts Before the Algorithm

AI in education is often discussed at the level of the model.
How accurate is it? How explainable is it? How personalised can the recommendations become?
Those questions matter. But in developmental assessment, there's an earlier question that matters just as much: what exactly does the data represent?
A recent UNOWA pilot study offers a useful example of why responsible AI in education has to start with that question, not skip past it.
What the Pilot Found
The study analysed data from 153 children aged 18–83 months across five functional domains: cognitive, social, physical, sensory, and daily living skills.
Four of those domains showed substantially stronger relationships with age. Sensory functioning didn't. Its correlation with age was much weaker, its associations with the other developmental domains were weak or near zero, and exploratory factor analysis placed sensory functioning strongly on a separate factor.
The paper's proposed interpretation matters here: sensory functioning may be better understood as a distinct regulatory construct rather than simply another age-graded developmental level.
Why does that matter for AI? Because an AI system can only interpret the structure it's given. If two different constructs get treated as if they mean the same thing, a more sophisticated algorithm doesn't fix the underlying measurement problem — it may just automate it.
Developmental Capability and Regulatory Access Are Different Questions
A developmental assessment helps us understand what a child can currently do. Can they solve a task? Communicate? Complete a physical action? Manage a daily living activity?
A sensory-regulatory profile asks a different question: under what conditions can the child access and use those abilities?
That distinction matters because similar observable difficulties can arise for different reasons. A child may struggle with a task because the task itself is beyond their current cognitive level. Another child may understand the task perfectly well but find it difficult to participate under sensory overload or during an unpredictable transition.
Those two situations shouldn't automatically lead to the same interpretation — which is exactly where the quality of the underlying assessment becomes critical for any AI system built on top of it.
Personalisation Is Only as Meaningful as the Construct Behind It
It's easy to describe AI-generated recommendations as personalised. But personalisation isn't simply producing different outputs for different users — the input has to carry a meaningful distinction in the first place.
The UNOWA pilot suggests sensory information may represent a different layer of functioning from the developmental domains analysed in the study. At the same time, the paper is clear that the current sensory item pool still needs substantial psychometric refinement: more differentiated sensory subscales, larger samples, test-retest reliability, external validation against established instruments, and renewed factor analysis after revision.
Both statements are true at once. The pilot offers a potentially useful distinction, and it tells us the measurement structure behind that distinction isn't finished.
AI Shouldn't Turn Preliminary Structure Into Artificial Certainty
The paper discusses AI-assisted interpretation as a future direction. It doesn't validate an AI system for diagnosis or automated recommendations, and that boundary matters.
Before AI can responsibly assist with interpretation, the constructs being interpreted need to be sufficiently clear and tested. Otherwise, the risk isn't just an inaccurate prediction — it's a system that produces a confident recommendation from a measurement model that's still too coarse. In that situation, the algorithm can look advanced while the uncertainty has simply moved upstream.
This is why responsible AI in education assessment begins before the AI layer, with questions like:
- Are we measuring one construct or several?
- Does the score behave as expected?
- Is it stable over time?
- Does it correspond with established measures?
- Does a global score preserve enough information for practical interpretation?
Those are measurement questions first. Only then do they become AI questions.
What the Pilot Changes
The most useful outcome of a pilot isn't a claim that the model is finished — it's a clearer map of what needs to be tested next.
Here, the study points in two directions at once. First, sensory functioning may need to be interpreted differently from age-graded developmental capability. Second, the sensory scale itself needs a more differentiated structure before stronger conclusions can be drawn.
That's a productive result. If future AI-assisted interpretation is going to support educators, specialists, or families, it shouldn't just generate more recommendations — it should generate recommendations from better-defined information.
The Responsible Sequence
The sequence shouldn't be: collect data → add AI → call it personalisation.
A more responsible sequence: define the construct → test the measurement → refine the structure → validate it → then explore AI-assisted interpretation.
The UNOWA pilot is one step in that process, not the end of it — and that's exactly why it matters. Good AI can't compensate for a poorly defined assessment model. Before asking what an algorithm can tell us about a child, we first need to be confident about what the assessment is actually measuring.
For the full data behind these findings, see our research explainer on the UNOWA pilot.
FAQ
Why does responsible AI in education depend on assessment quality? An AI system can only work with the structure of the data it's given. If a construct like sensory functioning is treated as equivalent to a developmental skill score, an algorithm built on top of that data inherits the same measurement error — it just presents it with more confidence.
Does this study validate AI for diagnosing sensory or developmental profiles? No. The paper discusses AI-assisted interpretation as a future direction, not a validated capability. The underlying sensory construct still needs refinement and external validation first.
What's the difference between a developmental score and a regulatory profile? A developmental score describes what a child can currently do. A regulatory profile describes the conditions under which the child can access and use that ability — a different question with a different set of implications for support.
What should come before AI-assisted interpretation in education? Construct definition, measurement testing, structural refinement, and external validation — in that order. AI-assisted interpretation is the last step, not the first.
Check out other articles
Explore the latest perspectives from our digital research team

Differentiated Support in Inclusive Education: Why a Score Is Only the Start
Differentiated support in inclusive education starts where scoring ends. Here's why assessment should answer "what's getting in the way," not just "what's wrong."

Sensory Processing in Child Development: What the UNOWA Pilot Found
A plain-language breakdown of the UNOWA pilot study on sensory processing in child development — the data, the factor analysis, and what still needs validation.

Streamlining RFP Responses for Educational Infrastructure
Discover effective strategies to streamline RFP responses for educational infrastructure projects in the EU, MENA, and CIS regions. Learn about best practices, innovative tools, and real-world examples to enhance your bid success and drive transformative educational reforms.

