Consider a district reviewing its first semester with a new AI platform. Teachers have activated their accounts, the professional development sessions are complete, and the usage report looks encouraging. Then someone asks how much planning time teachers have recovered, or whether students understand the material more deeply. The report has very little to say about either.
That is an uncomfortable position for a superintendent preparing a board update or a principal deciding what deserves another place in the budget. Schools need evidence that connects an investment to the work happening in classrooms. Account activity can help explain participation, but leaders also need to know whether the experience is worth continuing.
Measuring AI impact in schools begins with that connection. A useful review follows the time teachers spend, the quality of their work, and the learning students can demonstrate. The seven metrics below give leaders a way to examine those outcomes while keeping the evaluation manageable for the people doing the teaching.
Start with the problem the school wanted to solve
Before collecting data, return to the reason the initiative began. Perhaps teachers were spending too long adapting reading materials. Perhaps students needed more opportunities to practice explaining their reasoning. Those are different problems, and they require different evidence of progress.
A teacher planning tool should be evaluated for its effect on preparation time and instructional quality. A student tutoring tool needs closer examination of learning, independence, and access. Choose the measures that fit the intended use; an administrative pilot does not need to demonstrate gains in student test scores to justify its existence.
Write down the starting point before introducing the tool, then agree on what would count as a worthwhile improvement. The seven measures form a menu for that decision. Select a small number for the first review, and keep the others in view as the work develops.
1. Teacher time saved after checking and revision
Time is an understandable place to begin because teachers can describe exactly where it goes. Preparing a lesson may involve finding examples, adjusting the reading level, checking accuracy, and making sure the activity fits the students who will actually use it. An AI draft only helps if it reduces the total work involved.
A 2025 Gallup and Walton Family Foundation survey found that U.S. public school teachers who used AI at least weekly estimated saving 5.9 hours per week on average. Those were self-reported estimates, so they offer a reason to investigate local time savings rather than a target every school should expect to reach. [1]
Ask a small, representative group of teachers to record the time spent on a few comparable tasks before and during the pilot. Include prompting, fact-checking, revision, and any extra work created later. In a hypothetical example, a task that previously took 45 minutes and now takes 30 minutes from start to finish saves 15 minutes, even if the initial draft appeared almost instantly.
Track typical minutes saved per task and how much the results vary between teachers. Then discuss where that time went: additional student feedback, a less rushed planning period, or less work taken home. Those details help a principal understand whether the change is making the school day more manageable.
2. Teacher confidence supported by demonstrated skill
Time savings depend partly on whether teachers have enough support to use a tool well. A positive reaction at the end of a workshop tells you something about the session. Several weeks later, leaders need to understand what teachers can apply in their own subjects and where they still need help.
In a 2026 Gallup and Walton Family Foundation study, only 18% of U.S. public K–12 teachers reported receiving formal guidance from administrators on AI use. That finding makes clarity an important part of any school AI review: teachers need to understand which uses are appropriate and what review is expected of them. [2]
Combine a short confidence survey with a practical demonstration using a normal teaching task. A teacher might adapt an activity, explain the changes made to the AI output, and identify something they decided to leave out. Track the proportion who can complete the task to an agreed standard, alongside their confidence and requests for further support.
Make room for teachers who use AI selectively or decide it adds little value to a particular lesson. Their reasoning can reveal a poor fit, an unnecessary step, or a training need. Treating usage frequency as a performance expectation makes those conversations much harder to have honestly.
3. The quality of materials that reach students
A teacher who produces materials faster still needs confidence that those materials are worth teaching. A reading passage might be easier to understand but lose the vocabulary students need to learn. A polished explanation might contain an error that becomes apparent only when a student asks a good question.
An Education Endowment Foundation trial involving 259 teachers in 68 English secondary schools examined AI use for science lesson preparation. Teachers given ChatGPT and a supporting guide spent 31% less time on the preparation measured in the trial. An expert review found no apparent reduction in the quality of sampled resources, although that does not establish the quality of every AI-generated lesson or its effect on student learning. [3]
Schools can borrow the evaluation approach by reviewing a small sample of materials during existing team planning time. Agree on a few criteria before looking at the samples, so judgments are consistent:
- Check accuracy and curriculum fit. Review whether the content is correct, whether examples support the intended standard, and whether the material prepares students for the next part of the unit. Record the proportion of samples that meet the agreed expectations and the corrections teachers needed to make.
- Examine the thinking the task requires. Look at what students must explain, compare, justify, or apply. If an adaptation makes the language more accessible, check that it preserves the intellectual work students are supposed to do.
Use the findings to improve shared examples and training. A recurring weakness in generated questions, for instance, gives a coach a specific issue to address with teachers at the next planning meeting.
4. Learning students can demonstrate independently
The quality of a resource matters because of what students learn from it. A completed assignment offers only part of that evidence when AI has helped produce the response. Teachers also need opportunities to see whether students can explain the idea, recognize an error, or apply the skill in a different situation.
A 2025 study published in PNAS examined AI tutors with nearly 1,000 high school mathematics students in Turkey. Students using a basic GPT-4 interface performed better during supported practice but worse on subsequent unassisted assessments than students without AI access. A version designed with learning safeguards largely mitigated that negative effect. The study concerned a specific setting and short-term outcomes, which limits how broadly its results should be applied. [4]
For a student-facing pilot, build a brief independent check into ordinary instruction. After supported practice, students might solve a comparable problem, explain a choice aloud, or apply a concept to an unfamiliar example. Track performance over time using a consistent rubric, while preserving established accommodations and access supports.
Where feasible, compare progress with similar classes or tasks that did not use the tool. Note other changes, including extra instructional time or a new teaching approach, before attributing improvement to AI. The purpose is to understand what students have learned well enough to carry into the next lesson.
5. Students can evaluate AI responses responsibly
Independent subject knowledge also helps students judge AI output. A student who understands the topic has a better basis for questioning an explanation, checking a source, or recognizing that an answer leaves something important out. AI literacy should give students repeated practice making those decisions.
Assess this through short scenarios suited to the students' age and the tools their school permits. Use the same criteria over time and track the proportion who can explain a sound decision, rather than simply name a rule:
- Give students an answer that needs checking. Ask them to identify which claims require verification, consult reliable sources, and explain what they would revise. Credit the checking process and reasoning, including when students correctly conclude that there is not enough evidence to accept a claim.
- Ask students to choose an appropriate use of AI. Present a realistic assignment or privacy scenario and have them explain what help would be acceptable under school guidance. Their answer should show that they understand both the learning purpose and the boundaries around sharing personal information.
These activities can reveal gaps that a quiz about AI terminology would miss. They also create useful classroom discussions about why a fluent answer can feel convincing and what responsible judgment looks like when students are uncertain.
6. Access and benefits across different groups
As evidence begins to accumulate, look closely at who is represented in it. A pilot run entirely by enthusiastic volunteers can tell you what is possible under favorable conditions. It gives a less complete picture of what may happen when teachers with different experience, schedules, and support needs join the initiative.
Review participation and outcomes across relevant grade levels, subjects, and student groups, using only information the school is authorized to use and protecting small groups from identification. Track practical barriers as well: availability of training, compatible devices, accessibility features, language support, and whether an activity depends on a paid account at home.
An overall improvement can conceal uneven results. If one group is gaining time while another needs substantial troubleshooting, the next step may be targeted support or a different approach for that setting. Ask teachers and students what prevented them from benefiting, then check whether the change you make actually removes that barrier.
7. Total cost in relation to the benefit demonstrated
Once leaders understand the benefits and who receives them, they can make a more credible judgment about cost. The subscription price is only one part of that calculation. Include training time, substitute coverage where needed, technical support, accessibility work, and the staff time required to review and maintain the approach.
Choose a cost measure that matches the purpose of the initiative. A planning pilot might track total cost per participating teacher alongside verified time savings and resource quality. A student literacy program might compare its full cost with the number of students who meet an agreed performance standard, while considering other ways the school could achieve that goal.
Keep recovered staff time separate from cash savings. A teacher who gains 30 minutes for feedback has additional capacity, but that does not automatically reduce the district's payroll. Describe the benefit honestly so the budget discussion reflects what the school actually received.
Include the effort required to continue or change direction at renewal. A promising pilot may need more coaching before expansion, while a tool with limited benefit may warrant a smaller commitment. Clear evidence gives leaders a defensible basis for either decision.
Build an AI impact scorecard that teachers can live with
Collecting all seven measures every week would create its own workload problem. For an initial pilot, select two or three priority outcomes, assign one person to coordinate the review, and use information already available through planning meetings, student work, and brief staff check-ins. Set the review date to allow time for teachers to learn the approach and use it through a relevant teaching cycle.
A shared scorecard can remain simple while giving the leadership team enough context to act. For each selected metric, record the starting point, intended improvement, evidence collected, and the decision that evidence supports:
- Agree on success before the pilot begins. For a materials-preparation pilot, the team might seek a locally chosen reduction in preparation time while maintaining its existing quality standard. Record how both will be checked so enthusiasm or disappointment does not change the definition halfway through.
- Show variation and uncertainty in the review. Include the number of participants, a typical result, and the range of experiences, supported by a few examples. Where the evidence is incomplete, name the question that remains and decide whether a longer or better-supported pilot could answer it.
- Connect findings to a specific next step. Expand an approach when the relevant benefits are consistent and the school can support wider use. Adjust it when a solvable problem is limiting results, or end the pilot when the benefit does not justify the cost and effort.
Discuss the findings with the teachers involved before presenting a board summary. They can explain why an apparent improvement was harder to achieve in one class, or why a modest time saving mattered during a demanding part of the term. Their experience helps leaders interpret the numbers accurately.
Return to the leadership meeting where the usage report left the central questions unanswered. With a focused review, the conversation becomes more specific: teachers can describe the time recovered, students can demonstrate the skills learned, and leaders can explain what deserves further investment. Families and board members have evidence they can examine, including the limits of what the school knows so far.
At TomoClub, we work with schools on AI literacy, teacher professional development, and implementation support. If your school is planning an initiative or reviewing one already underway, connect with our team to discuss the problem you want to solve and the evidence that would make progress meaningful in your classrooms.
Sources
- Gallup and Walton Family Foundation, June 24, 2025. Survey of U.S. public K–12 teachers reporting AI use and estimated time savings.
- Gallup and Walton Family Foundation, May 26, 2026. Research on formal and informal guidance for teachers using AI.
- Education Endowment Foundation, ChatGPT in lesson preparation. Teacher Choices trial independently evaluated by NFER.
- Bastani and colleagues, PNAS, 2025. Generative AI without guardrails can harm learning: Evidence from high school mathematics.