An Eight-Week Pilot Protocol for Any Dental AI Tool
Most dental AI pilots end in a purchase. Not because the software earned it, but because nobody wrote down, in advance, what failure would look like.
You know how it goes. A rep demos something at the BDA conference or on a Teams call. The images look sharp, the caries detection boxes land neatly on the interproximals, and someone on the team says “let’s just try it for a bit.” Three months later there’s a monthly direct debit and a login nobody at reception uses. When you ask whether it worked, you get impressions rather than numbers: “the associates seem to like it”, “patients think it’s clever”, “I think it’s caught a few things.”
That’s not a pilot. That’s an unusually long sales process where you paid for the last leg.
This piece gives you a protocol you can run on any dental AI tool: radiograph reading (Pearl Second Opinion, Overjet, VideaHealth, Dentalyze), patient triage and recall chasers, or front-desk admin like automated notes and phone answering. Eight weeks, three measured phases, and one pre-agreed kill criterion written before the trial licence is activated. If you want the wider decision framework this sits inside, the Choosing and Integrating Tools pillar covers procurement, DSP integration and IG questions in more depth. This article is only about the trial itself.
Why pilots without a kill criterion always convert
Here’s the mechanic. Once a tool is in the surgery, three things accumulate that have nothing to do with whether it works.
Sunk effort. Somebody spent a Wednesday afternoon getting it talking to your imaging software. That afternoon becomes an argument for keeping it.
Social cost. The associate who championed it will be embarrassed if it’s dropped. The rep has been friendly and helpful and answered emails at 9pm. Cancelling now feels like a personal rejection, so the decision gets deferred rather than made.
Absence of counter-evidence. Nobody measured anything before switching on, so there’s no baseline to compare against. In the absence of numbers, the loudest impression wins, and the loudest impression is always from whoever liked it most.
A kill criterion breaks all three. It converts “should we keep this?” into “did the number we agreed on get hit?” You are no longer disappointing your rep. You are reading a threshold you set eight weeks ago, when nobody had any skin in the game.
Phase 1 (weeks 1–2): baseline with the tool switched off
Do not install anything yet. If you have already installed it, turn the AI overlay off for two weeks and collect this data anyway.
You need a baseline because AI vendors quote uplift figures against a comparison group you’ve never met. Pearl’s FDA clearance material and Overjet’s published claims both describe sensitivity improvements over unaided reads, and those studies are real, but they are not your practice. Your baseline is what you’re actually buying an improvement on.
For radiograph AI, pull these from your last completed quarter, per clinician:
| Metric | Where to get it | Example (3-surgery mixed practice) |
|---|---|---|
| Bitewings taken per 100 exams | Imaging software report | 62 |
| Restorations planned per 100 bitewing sets | Treatment plan export | 18 |
| Periapicals retaken (image quality) | Imaging software audit log | 9% |
| Mean time from radiograph to plan discussion | Stopwatch, 20 consecutive cases | 4m 10s |
| Second-opinion requests between associates | Ask them to tally for 2 weeks | 3 in 2 weeks |
For triage and front desk, the baseline is operational:
| Metric | Example baseline |
|---|---|
| Unanswered inbound calls, 8am–6pm | 71 per week |
| Mean time to first response on web enquiries | 6h 40m |
| FTAs as % of booked appointments | 8.1% |
| Reception overtime hours per week | 5.5 |
| Emergency slots filled with genuine emergencies | 64% |
Two weeks of measuring with the AI off is the single most-skipped step in dental AI pilots and the only reason most of them are unfalsifiable. It costs you a shared spreadsheet and about ten minutes a day at reception.
Phase 2 (weeks 3–6): the tool runs, shadowed
Switch it on. Critically, for diagnostic tools, run it in shadow mode for the first two weeks: the clinician reads the radiograph and records their findings, then reveals the AI overlay and records whether they changed anything. If your software can’t sequence it that way, use a simple paper tally: finding, AI agreed / AI flagged extra / AI missed, and whether you altered the plan.
This distinguishes the two things that matter and get conflated constantly.
Agreement is the tool telling you what you already saw. Pleasant, reassuring, worth roughly nothing. If a tool agrees with you on 94% of surfaces, you have paid £400 a month for confirmation.
Actionable disagreement is the tool flagging something you then verified and acted on, or you overruling it and being right. Both are useful. Only the first is what the vendor is selling.
A worked example from a four-surgery mixed practice in the Midlands running Pearl on a trial licence, 340 bitewing sets over four weeks:
AI flags on surfaces the clinician had not marked: 211
Reviewed and agreed by clinician (plan changed): 24
Reviewed and dismissed (enamel-only / artefact): 187
Clinician findings the AI did not flag: 38
Confirmed on re-examination: 31
Clinician revised own finding after AI disagreement: 7
Net plan changes attributable to the tool: 24 additions, 7 removals
False-positive review burden: 187 dismissals
Mean added time per bitewing set (measured): 38 seconds
Twenty-four genuine additions across 340 sets is roughly one in fourteen patients getting a lesion caught earlier. That’s clinically meaningful. It also cost 187 dismissals and about 3.6 hours of cumulative clinician time across the month. Whether that trade is worth £350–£600 a month is a judgement, but it’s now a judgement about two real numbers rather than a feeling about a demo.
By week 5, drop shadow mode and use the tool the way you’d use it in practice. Keep tallying. The interesting shift in week 5 and 6 is behavioural: watch whether clinicians start pre-emptively deferring to the overlay rather than reading first. If an associate tells you they’ve stopped looking as carefully because “Pearl will catch it”, that is a finding, and a serious one. Write it down.
The kill criterion: write it before week 1
This is the whole protocol, really. Everything else is measurement infrastructure to make this line enforceable.
The kill criterion is a single sentence with a number and a date, signed off by the principal, and stored somewhere other than the champion’s inbox. Put it in the practice meeting minutes.
Good ones look like this.
Radiograph AI: “If by 24 November the tool has produced fewer than 15 clinician-confirmed findings we would otherwise have missed across at least 250 bitewing sets, or has added more than 45 seconds mean read time per set, we cancel and do not renew the conversation for 12 months.”
Front-desk phone AI: “If unanswered calls have not fallen from 71 to below 25 per week by week 6, and reception overtime has not fallen by at least 3 hours per week, we cancel. Patient complaints about speaking to a machine above 3 in the pilot period also triggers cancellation regardless of the call numbers.”
Triage / recall chaser: “If FTA rate has not moved below 6.5% and the tool has generated more than 5 messages requiring an apology to a patient, we cancel.”
Notice the shape. A metric that was baselined in Phase 1, a threshold, a deadline, and a consequence. Notice also the second clause in two of them: a safety or reputational tripwire that overrides the efficiency win. A phone AI that answers every call while irritating patients is not a success, and if you don’t write that down beforehand, the call-answering graph will win the argument.
Three rules keep the criterion honest. Set the threshold before you see any pilot data. Give one named person authority to declare the kill, and make it someone who didn’t source the tool. Do not allow “let’s extend the trial to gather more data” as an outcome: an extension is a purchase with extra steps, and every vendor will happily grant you one.
Phase 3 (weeks 7–8): decide, then write the reversal plan
Week 7 is reconciliation. Put the Phase 1 baseline next to the Phase 2 numbers on one page and read the criterion aloud in the practice meeting. It either hit or it didn’t.
If it didn’t hit, cancel that week. Not next month. Vendors know that the gap between “we’ve decided against it” and “we’ve actually turned it off” is where renewals live, and 30-day notice clauses are typically written to exploit exactly that gap. Send the cancellation email while the meeting is still happening.
If it hit, week 8 is not for celebrating, it’s for writing down how you’d get out. Three things: where your data sits and how you’d export it (a DPIA question as much as a commercial one, and your processor agreement should already cover it), which workflows now depend on the tool and what the manual fallback is if it’s down for a day, and a review date twelve months out with the same metrics re-measured. Tools decay. Vendors get acquired, models get retrained, the version that impressed you in week 4 is not the version running next April.
One more thing worth doing in week 8: ask the whole team, separately and in writing, whether they’d notice if it vanished on Monday. Reception, nurses, associates, hygienists. It’s a cheap question and it has caught at least one practice I know of paying £4,800 a year for a caries detection licence that two of four associates had quietly stopped enabling.
The awkward version of all this is that a properly run pilot will sometimes tell you a genuinely good tool isn’t good enough for your practice yet, at your volumes, at that price. Twenty-four extra lesions across 340 patients is a real clinical win and might still not clear a threshold you set honestly. That’s allowed. The point of the number is that it decides for you on a day when everyone in the room, including you, has already started wanting to say yes.