A twelve-minute vendor demo can make a weak instructional idea look like a breakthrough. A chatbot produces a polished debate rubric, three reading levels, and instant feedback while adults nod. None of that tells a principal whether students learned to argue, read closely, or revise their thinking. A principal AI guide should begin with a firm rule: evaluate an AI tool as an instructional intervention, not as a software purchase. Approve it only when a bounded trial shows that it advances a named learning goal without quietly transferring judgment, privacy, or repair work to teachers and students.
This standard does not require every school to run a randomized trial. It requires leaders to ask a better question than whether the product feels impressive. The relevant unit is the learning task: what students must know, do, explain, or remember after the tool has been used.
Start with a learning claim, not a feature list
Most procurement conversations begin with capabilities. Can the tool generate questions? Can it differentiate text? Can it score writing? Those are software functions, not reasons for students to learn.
When evaluating EdTech, borrow a simple idea from evidence-centered design, associated with Mark Mislevy and colleagues: connect a claim about learning to observable evidence and to the task that should produce it. A principal and teacher can put that connection on one page:
- Claim: Students will explain a historical cause using relevant evidence.
- Evidence: Students select evidence, make an inference, and defend the connection in their own words.
- Tool role: The AI asks a counter-question or points out an unsupported step.
- Boundary: It does not select the evidence or write the explanation.
That page exposes a common problem. If the tool supplies the interpretation, evidence selection, and polished prose, the final essay may look stronger while showing very little student reasoning. The product has improved performance in the moment, but the learning claim has not been demonstrated.
Research on retrieval practice and Robert Bjork’s work on desirable difficulties offer a useful warning here: ease during a lesson can hide weak retention. A tool that gives the answer quickly may be less useful than one that offers a partial clue and requires a student to retrieve, choose, or revise. The most capable AI tool is sometimes the wrong tool for the first version of a lesson because it removes the work that the lesson was meant to make visible.
A Year 8 science team, for example, should score causal reasoning separately from grammar and formatting. If students use AI to polish an explanation, they still need to identify the variable, select evidence, and explain the mechanism without the tool doing those steps for them.
The strongest case for moving fast deserves respect
The strongest argument against this cautious approach is practical. Schools cannot wait for perfect research on every new tool. Teachers are already experimenting, often without clear guidance. A long approval process may push use into private corners of the school, where leaders cannot support teachers or see problems. A blanket ban can also deny students useful practice and leave access to families with more money, better devices, or more time.
There are legitimate low-stakes uses. A teacher may save time by asking for a first draft of a parent newsletter, a set of examples, or alternative explanations for a difficult concept. Students may benefit from language support, structured questioning, or immediate opportunities to revise. A principal should not demand evidence of higher standardized test scores before allowing a teacher to test a planning aid.
That argument changes the pace and scope of evaluation, not the need for judgment. Teacher-only drafting, student brainstorming, automated feedback, and high-stakes assessment should not move through the same approval lane. The closer a tool gets to grading, placement, discipline, or a student’s independent demonstration of knowledge, the stronger the evidence and adult oversight should be.
Fast permission for a small, reversible experiment is sensible. Fast school-wide adoption based on a polished demonstration is not. The difference is whether the school has stated what it is trying to improve and what would make it stop.
Ask for evidence around the tool, not a dazzling demo
A product demonstration is a sample of performance under favorable conditions. It is not evidence that the tool works in your curriculum, with your students, under normal teacher workload. A serious evaluation asks for artifacts that make the hidden work visible.
What should a principal request?
Before or during a pilot, ask for four things. First, collect raw outputs from several prompts written by local teachers, including an ambiguous prompt and a weak student response. Edited showcase examples conceal the checking and rewriting that make the result usable. Second, attach the actual student task and the school’s rubric. A generic claim about personalization means little if nobody can say which rubric criterion should change. Third, record teacher time for setup, checking, correction, and follow-up—not merely the minutes saved by generating a first draft. Finally, state what happens if the school ends the trial: how teachers retain their prompts, rubrics, and curriculum materials, and how student records are handled.
This is where evaluating EdTech becomes an instructional exercise rather than a procurement exercise. Have teachers score a small sample of student work against the same rubric used for the non-AI version. If possible, ask someone who did not design the lesson to score some samples without knowing which version produced them. The exercise will not prove causation, but it can reveal whether the supposed gain is reasoning, accuracy, formatting, or simple completion speed.
Generated material also needs an adult review path before students see it. During a pilot, that may mean a teacher checks every question, example, or feedback prompt. If the school cannot staff that check, the proposed use is too broad for the moment. This is a real cost, not an administrative footnote.
A student satisfaction survey can help identify confusing instructions or poor accessibility. It cannot establish that students learned more. Ask students what they decided, changed, or rejected while using the tool. Their answers are more useful than a single rating about whether the experience was fun.
Pilot with a stop rule
A pilot should have a stop date before it has a start date. For a first test, choose one learning objective, one grade team or two volunteer teachers, and three lessons across two or three weeks. That is small enough to observe closely and large enough to expose routine problems.
Use a baseline task before the AI-supported version and a comparable task afterward. Keep the rubric stable. If the classes are not comparable, say so; a local pilot is a decision aid, not proof that the tool caused a gain.
At minimum, collect four signals:
- Student work on the target skill, with polish scored separately from reasoning.
- A short delayed check one or two weeks later, without the tool, to test retention or transfer.
- Teacher time spent preparing, checking, repairing, and explaining the tool.
- A brief student account of what the AI contributed and what the student still had to decide.
Set the threshold before reviewing results. A school might decide that a teacher-only planning tool must save at least ten minutes per lesson without lowering quality, while a student-facing feedback tool must show improvement on the target rubric criterion without making students less able to explain their choices. The exact number is a local judgment. The important point is that leaders do not move the goalposts after seeing an attractive demo or a handful of excellent outputs.
This approach fails when leaders treat a tiny pilot as causal proof, or when they demand the same evidence standard for every use. A scheduling assistant should be judged mainly on reliability and staff time; a tool that recommends intervention groups requires much stronger evidence about accuracy, fairness, and consequences. The evidence base for many classroom AI uses is still thin, so a successful pilot should earn another test, not automatic permanence.
School AI adoption needs an owner and an exit ramp
Adoption is not complete when accounts are created. Someone must own the instructional purpose, the review process, and the decision to pause. The NIST AI Risk Management Framework offers a useful sequence for school leaders: govern, map, measure, and manage.
In practice, govern means naming the responsible leader and the teacher team, while stating the approved use. Map means identifying who is affected, what data enters the system, where students might rely on it too heavily, and what can go wrong. Measure means reviewing learning evidence, staff workload, access, and student understanding. Manage means setting a review date and a clear route to revise, restrict, or retire the use.
That record should also protect teacher authorship. If teachers create prompts, rubrics, examples, or activity designs during a pilot, the school should keep copies in a common format and make clear who may revise them. Curriculum should not become unusable because a license ends or a vendor changes a feature.
Privacy belongs here as a pass-or-fail gate, not as one more attractive score in a procurement spreadsheet. Before student-facing use, leaders should know whether accounts are required, what information is retained, how deletion works, who can see activity logs, and whether the use is appropriate for the students’ ages. If those answers are unclear, the instructional promise does not outweigh the unresolved risk.
Before Friday, take one proposed tool and write five sentences: the learning goal, the student behavior that will show it, the AI’s permitted role, the evidence to collect, and the stop rule. If the team cannot complete that page without falling back on words such as engagement or personalization, the school is not ready to purchase. It may be ready to run a smaller experiment—and that is the more responsible next step.

