edcircuit
Share Your Voice on edCircuit
Promotional banner with a book cover on the left titled 'Purposeful Technology, Powerful Learning: You Can't Ban Possibility' and a blue text panel on the right reading '13 districts. Real classrooms. Documented stories of technology used with intention. Learn More: stayh... com' with the CoSN logo at the bottom.
Home InnovationArtificial Intelligence AI and Student Learning: How Schools Should Measure It
10 minutes read

AI and Student Learning: How Schools Should Measure It

Why adoption rates and time savings are not enough to prove that AI is improving education

AI and student learning should be evaluated through achievement, independence, equity, teacher workload, user feedback, and unintended consequences.

AI and student learning are supposed to be the focus of the school board presentation.

Instead, the first slides show how many accounts have been activated, how often the new platform is used, and how many lessons, quizzes, summaries, and tutoring sessions it has generated. Another slide estimates that teachers have saved hundreds of hours.

The numbers look impressive.

Then a board member asks the question that is missing from the presentation:

“Are students learning more?”

The room goes quiet.

The district knows how many people logged in and which features they used. It may even know how much time teachers believe they saved. What it cannot clearly explain is whether students understand more, think more deeply, work more independently, or receive stronger instruction because the technology is there.

That is the measurement problem schools now face.

AI activity is easy to count. Educational value is harder to establish.

As districts move from experimentation to long-term implementation, enthusiasm must be matched by evidence. The central question cannot be whether AI is being used. It must be whether the technology is improving teaching and learning—and at what cost.

Start With the Question, Not the Dashboard

Usage data has a natural appeal. It is immediate, precise, and easy to place in a report.

A district can count active users, prompts, completed tutoring sessions, and minutes spent on a platform. Those numbers help leaders understand whether accounts work, whether training reached the intended staff, and whether people are using the features the district purchased.

They do not demonstrate learning.

A student can complete fifty tutoring sessions without mastering the targeted skill. A teacher can generate dozens of lesson plans that still require significant correction. A writing assistant can produce more polished essays while students become less capable of drafting independently.

According to National Center for Education Statistics data released in 2025, 73 percent of public schools reported that at least some teachers were using AI for lesson planning, administration, instructional materials, assessments, grading, or feedback. Yet only 38 percent of school leaders agreed that integrating AI would lead to better educational outcomes for students.

That gap matters. Schools are already using the technology, but its educational value remains unsettled.

A 2026 resource from Regional Educational Laboratory Central similarly notes that research has not established which teacher uses of AI translate into improved student learning outcomes.

Districts should not interpret that uncertainty as a reason to avoid every AI tool. They should treat it as a reason to define success before implementation.

Suppose a district adopts an AI tutoring system for middle school mathematics. “Improving math” is too broad to evaluate. Leaders need to identify which students will use the tool, which skills it is expected to strengthen, how teachers will use the information, and what improvement should be visible.

Will students solve more multistep problems correctly? Can they explain their reasoning? Will they retain the skill after the tutoring support is removed? Can they apply what they learned to an unfamiliar problem?

Those questions point toward evidence of learning. Account activations do not.

Without clear goals, success is easily redefined after the fact. A platform purchased to improve achievement may later be defended because students enjoyed using it. A tool intended to save time may be called successful because it produced more materials.

Enjoyment and productivity may be worthwhile. They are not interchangeable with learning.

Efficiency Matters—But So Does What Students Can Do

Teacher time is valuable.

If AI can reduce repetitive administrative work, provide a usable first draft, or make routine communication more efficient, districts should measure that benefit. Time returned to instruction, feedback, collaboration, or individual student support has real value.

But a tool that saves a teacher twenty minutes has not automatically improved a student’s education.

Leaders need to understand what happened during those twenty minutes. Did the teacher provide more feedback? Meet with a struggling student? Strengthen the lesson? Or did the time disappear into correcting errors the AI introduced elsewhere?

Efficiency measurements should include time added as well as time saved.

A quiz may be generated in five minutes but require twenty minutes of fact-checking and revision. A counselor may receive a faster summary of student information but need additional time to verify what was omitted. A translated family message may appear instantly, yet someone still needs to confirm that it is accurate and appropriate.

During a pilot, participating educators can track a small sample of tasks before and after using the tool. How long did the work take? How much correction was required? Was the final result stronger, weaker, or essentially the same?

The same care is needed when examining student work.

AI can improve the appearance of an assignment without strengthening the thinking behind it. An essay may sound more polished. A science response may use advanced vocabulary. A presentation may be better organized. None of those changes proves that the student understands more.

For AI-supported writing, teachers might compare a student’s ability to plan, draft, revise, and explain key choices. An independent writing sample can reveal whether those skills transfer when the technology is unavailable.

For an AI tutor, districts should look beyond performance inside the platform. Can students solve a new problem during class? Can they explain why the solution works? Do they remember the skill several weeks later?

Student independence deserves attention. If performance improves only when the tool is present, the system may be supporting completion rather than durable learning. That support may still have value, but schools should be honest about what it accomplishes.

Ask Who Benefits

A districtwide average can hide very different experiences.

An AI platform might improve outcomes overall while producing weaker results for English learners, students with disabilities, students using certain dialects, or children whose backgrounds are poorly represented in its training data.

Access may create another divide. Students with reliable devices, strong internet connections, and quiet places to work can use a platform differently from classmates who have access only during school hours. Students who already know how to ask detailed questions may receive stronger responses than those who need more guidance.

Districts should review participation, achievement, error reports, accessibility concerns, and teacher overrides across relevant student groups while protecting individual privacy.

They should also ask whether the tool expands opportunity or simply serves students who were already succeeding.

Imagine an AI-supported advising platform that increases the number of students receiving college and career information. That sounds promising. But are first-generation students receiving the same quality of recommendations? Are students with disabilities being directed toward a narrower set of options? Can multilingual families understand and question the guidance?

The average outcome will not answer those questions.

Neither will a vendor dashboard. Teachers, students, and families must be part of the evaluation.

A teacher may notice that a platform performs well in one subject but repeatedly generates inaccurate material in another. Students may report that an AI tutor gives away answers too quickly, making it easy to avoid productive struggle. Families may be uncomfortable with how a system describes or classifies their child.

Those experiences are evidence.

The NIST AI Risk Management Framework’s measurement guidance emphasizes that AI benefits and risks emerge from the interaction between technology, people, and the context in which a system is used. NIST recommends testing whether systems function as claimed, tracking errors and negative effects, comparing performance before and after deployment, and involving users in evaluation.

For schools, that means looking inside classrooms rather than relying entirely on platform analytics.

Measure Value and Harm at the Same Time

No single number can capture the educational impact of AI.

A useful district evaluation brings together several kinds of evidence: student performance, the quality of student reasoning, ability to complete independent tasks, teacher workload, differences among student groups, reported errors, accessibility concerns, privacy incidents, and feedback from the people using the system.

The evidence should also be compared with a meaningful starting point. What happened before the platform was introduced? What happens in similar classrooms that are not using it? Could an existing intervention address the same need at a lower cost or with less risk?

Perfect research conditions are rarely available in a school district. Thoughtful comparisons are still possible.

A limited pilot across several classrooms can provide useful information when the goal is clearly defined and starting conditions are documented. The district is not trying to publish a national study. It is trying to make a responsible local decision.

The review should begin early.

During the first month, leaders can focus on access, technical problems, training needs, privacy questions, and whether the tool functions as promised. After several months, they can examine teacher workload, student participation, output quality, accessibility, and early evidence of learning.

By the end of a semester or school year, the district should be able to compare academic outcomes, student independence, equity, cost, teacher practice, and unintended consequences with the original goals.

Harms and near misses belong in that review.

Did the platform produce inaccurate instructional content? Were students incorrectly flagged for misconduct? Did teachers enter information that should not have been shared? Were inaccessible materials generated? Did the system underestimate certain students and recommend lower-level work?

A teacher who catches a dangerous science error before a lesson has identified a weakness worth documenting. So has a counselor who finds that an automated summary omitted important context.

The NIST guidance on managing AI risk recommends continued monitoring because a system’s performance can change after deployment. It also calls for feedback, incident response, appeals, overrides, and the ability to discontinue systems that exceed an organization’s tolerance for risk.

Approval should begin the monitoring process, not end it.

Let the Evidence Change the Decision

An evaluation has little value if the district has already decided to renew the contract.

Leaders should determine in advance what would justify expansion, what would require modification, and what would cause the district to stop using the system.

Some tools will deserve broader use. Others may work only in certain subjects or grade levels. A district might retain one feature, disable another, strengthen its training, or revise the rules governing student use.

Sometimes the evidence will show that the platform should be removed.

Ending a pilot is not a failure when the pilot was designed to answer a question. Continuing an ineffective system because the district has already invested money and time is the more expensive mistake.

School boards and communities deserve an honest account of what was learned. Leaders should explain why the tool was introduced, who used it, which outcomes were examined, what improved, what did not, and what concerns emerged. Limitations in the evidence should be acknowledged rather than hidden behind early statistics.

The U.S. Department of Education’s guidance on AI in schools emphasizes responsible adoption, privacy, and engagement with affected stakeholders, especially parents.

Transparent evaluation is part of that responsibility.

Families do not need to be told that a platform is transformational. They need to understand what it does, why the district uses it, what evidence supports that use, how students are protected, and how concerns can be raised.

Teachers should see that their feedback influenced the decision. Students should understand how the technology affects their learning and what they can do when something goes wrong.

Trust grows when districts are willing to discuss uncertainty rather than hide it behind a dashboard.

Measure What Matters

AI may help educators work more efficiently, give students faster feedback, broaden access to support, and create learning experiences that were previously difficult to provide.

Those possibilities are worth exploring.

They are not self-proving.

A tool should earn its place in a school by serving a clearly defined educational purpose, producing benefits that can be observed, and operating within boundaries the community understands.

The school board presentation should include more than accounts activated, prompts submitted, and hours reportedly saved.

It should show what students understand now that they did not understand before.

The number of times a tool is used tells a district whether it has been adopted.

It does not tell leaders whether it deserves to remain.

Subscribe to edCircuit to stay up to date on all of our shows, podcasts, news, and thought leadership articles.


Editor’s note: edCircuit is vendor-neutral and does not endorse products or services. Any mention of a specific solution is included for contextual and informational purposes only.

  • edCircuit is a mission-based organization entirely focused on the K-20 EdTech Industry and emPowering the voices that can provide guidance and expertise in facilitating the appropriate usage of digital technology in education. Our goal is to elevate the voices of today’s innovative thought leaders and edtech experts. Subscribe to receive notifications in your inbox

    View all posts
Blue hero banner promoting a CoSN resource on screen time policies for district leaders, with a booklet image and the URL CoSN.org/screentimepolicy (From Screen Time Mandates to Purposeful Technology Policies and Guidance).

Join Thousands of Other Subscribers

This field is for validation purposes and should be left unchanged.

Participate in the COmmunity

Blue hero banner promoting a CoSN resource on screen time policies for district leaders, with a booklet image and the URL CoSN.org/screentimepolicy (From Screen Time Mandates to Purposeful Technology Policies and Guidance).
Science Safety - Safer Labs, Safer STEM, Safer CTE, Safer Arts, Safer Cyber

Use EdCircuit as a Resource

Would you like to use an EdCircuit article as a resource. We encourage you to link back directly to the url of the article and give EdCircuit or the Author credit.

MORE FROM EDCIRCUIT

-
00:00
00:00
Update Required Flash plugin
-
00:00
00:00