← All articles

How Should Schools Pilot an AI Tool? A Six-Week Plan

A week-by-week plan for piloting an AI tool with real students and teachers: what to decide first, what to measure and how to make the call at the end.

A demo tells you what a tool can do. A pilot tells you what it does in your school, with your students, on an ordinary Tuesday. Six weeks is long enough to get past the novelty and short enough to keep a group's attention, and it is the length this plan assumes. It goes with our procurement scorecard, and it works for tutoring tools, practice tools and teacher-facing tools alike.

The short answer to how a school should pilot an AI tool is to decide the question, the group, the comparison and the stop rule before you start, measure use and learning as well as satisfaction, watch closely in week three when the novelty fades, and make the final call against criteria you agreed in advance.

What should you decide before the pilot starts?

Four decisions shape everything else. The first is the question. Write one sentence in the form "We are piloting [tool] to find out whether [specific change]," for example whether students who use it for fractions practice do better on an unaided check than similar students who do not. A pilot without a question tends to end with "people seemed to like it," which is not a basis for a decision.

The second is the group. Choose one or two grade levels or subjects and at least two teachers who are willing. Enthusiasm helps, but include at least one sceptic, because their feedback will be the most useful. The third is the comparison. Without something to compare against, you cannot say whether the tool made a difference, and the options are another class, the same class in a previous year, or a group that starts later. The fourth is the stop rule, meaning what would make you stop early, such as a privacy concern, an accuracy problem or a serious safety incident. Agreeing this in advance makes it much easier to act if it happens.

What should you measure?

Measure more than satisfaction. Students can like a tool that does not help them learn, so pair opinion with use and with an outcome.

Area Measure How
Use Share of invited students who use it each week Vendor dashboard or teacher tally
Depth Sessions per student, and length Vendor data
Learning Result on a short unaided check Same test before and after
Teacher experience Time saved or lost, and confidence Short weekly survey
Student experience Whether they find it useful Two-question survey
Safety and accuracy Incidents, errors found Shared log

Track usage from the first week, because access is not the same as use. Khan Academy has said that only about 15 percent of students with access to Khanmigo regularly use it (EdTech Innovation Hub), and a randomised trial in Tennessee found students mostly stopped using it when it asked questions instead of giving answers (Hechinger Report). A tool that is well designed but unused teaches nobody anything.

What happens in each of the six weeks?

Week 0 is set-up. Sign the agreement only after you have read the data terms, and confirm what student information the tool collects and who can see it. Give teachers an hour to try the tool themselves before students see it, and write down the baseline: the unaided check, and current usage of any similar tool.

Week 1 is introduction. Show students how to use the tool and what you expect, which is to attempt first and then ask. Run the baseline check, and ask teachers to log anything odd, however small. Small oddities in the first week are often the early sign of a larger problem.

Week 2 is settling in. Check who is actually using it. If usage is low, ask why before changing anything, because the answer decides what you do next. Ask each teacher one question, what has surprised you, and review the incident log.

Week 3 is the dip. Novelty wears off around now, and usage often drops. This is the most valuable week to observe. Talk to a few students about why they use the tool or do not, and note what would make them return.

Week 4 is looking at the work. Read a sample of student sessions, with permission and following your privacy rules. Is the tool asking questions and giving hints, or handing over answers? Compare what students say with what you see.

Week 5 is checking learning. Run the unaided check again, with a similar but not identical version, and collect the teacher and student surveys.

Week 6 is the decision. Meet with the pilot group and work through five questions. Did students use the tool, and keep using it? Did they learn more than the comparison group on an unaided check? Did teachers gain something, in time or insight, that outweighed the cost? Were there safety, privacy or accuracy concerns? And what would it cost to run at scale, including training?

How should you make the final call?

Use four outcomes so that you do not default to "let's keep going." Adopt if there is clear evidence of use and learning and no unresolved concerns. Adopt with conditions if the tool is promising but needs changes to how it is used or to the contract. Extend the pilot if you do not have enough data, and name a specific question for another term. Stop if use is low, there is no learning benefit, or a concern stays unresolved. A clear "stop" is a legitimate result of a pilot, not a failure of it.

How do you get teachers on board?

Teachers are not a single audience. In an EdChoice survey, 55% of teachers opposed using AI in the classroom and 38% supported it, and the same survey found that 72% agreed it is important to help students build the critical thinking skills to use AI appropriately (EdChoice). Scepticism is common, so plan for it.

Start with a problem teachers already have, such as time spent on planning or feedback, and not with the technology. Invite sceptics into the pilot and treat their concerns as design input. Train before you launch, since an Education Week survey found 58% of educators had received no AI training (Chalkbeat). Share honest results, including what did not work, and give teachers a say in which tools continue after the pilot.

What are the common pilot mistakes?

The most frequent is piloting only with enthusiasts, which tells you what works for them and not for everyone. Measuring only satisfaction is a close second. Running a pilot with no comparison means results cannot be attributed to the tool. Some schools skip the privacy check because the trial is free, but student data is still involved. And many pilots decide on the last day, when the criteria should have been agreed in week zero.

How should you communicate about the pilot?

Tell families and staff about the pilot before it starts, what data is involved and how to raise a concern. Share a short summary of the result afterwards, whichever way it goes. A pilot that ends in "no" and is explained well builds more trust than one that ends in "yes" and is not.

Frequently asked questions

How many students do we need? Enough that a difference would show. For a first pilot, two or three classes with a comparison group is usually workable, and the aim is a clear signal on use and learning, not statistical proof.

Can we pilot more than one tool at once? It is possible, but each needs its own group and comparison, and the workload doubles. Most schools do better running one pilot well.

What if the vendor supplies its own evaluation? Use it as information, and still collect your own data on use and on an unaided check. Vendor evidence rarely matches your students and setting.

Who owns the decision? Decide before the pilot starts, and include the teachers who ran it. A decision made by someone who saw only a summary tends to be revisited.

Related guides