Sensitivity, Specificity and Why Your AI Flags Caries That Isn’t There
You bought the software because the vendor showed you a number. Probably something like 95% sensitivity for proximal caries, benchmarked against a consensus panel of three examiners on a curated set of bitewings. It’s a real number. The study behind it is probably real too. And then you switched it on for a Tuesday morning of six-month recalls and it started drawing boxes on every third contact point, including several you would stake your registration on being sound.
Nothing has gone wrong. The tool is doing exactly what it was tuned to do. The problem is that it was tuned on a population that looks nothing like yours, and nobody explained to you that the sensitivity number and the number of false alarms you get on a Tuesday are two ends of the same lever.
The lever nobody shows you at the demo
Every caries detection model outputs a probability, not a diagnosis. Internally it says something like “0.61 confidence that this region contains a lesion.” A threshold decides what you actually see. Set the threshold at 0.3 and the box appears. Set it at 0.7 and it doesn’t. That single number is the difference between a tool that catches everything and misses nothing, and a tool that catches everything and flags everything.
Sensitivity is the proportion of genuine lesions the model flags. Specificity is the proportion of genuinely sound surfaces it leaves alone. You cannot raise one without lowering the other on a given model: you can only slide the threshold. Vendors publish the point on that curve that makes the marketing deck sing, which is almost always the high-sensitivity end, because “misses 1 in 20 lesions” reads worse to a buyer than “occasionally over-calls.”
There’s a third number that matters more than either, and it barely appears in the literature dentists get shown: positive predictive value. When the software draws a box, what is the chance there’s actually a lesion in it? That’s the only question you’re answering at the chairside. And PPV depends on something the model knows nothing about: how much disease is in the mouths walking through your door.
The arithmetic, done properly
Take a model at 95% sensitivity and 85% specificity. Those are defensible real-world figures. Published evaluations of the commercial systems cluster in this region: Overjet, Pearl’s Second Opinion, VideaHealth, Denti.AI, dentalXrai. Sensitivity for proximal caries in the high 80s to mid 90s, specificity typically 80 to 90 depending on where the vendor has parked the default threshold and whether enamel-only lesions count as positives.
Now apply it to a bitewing set.
A four-bitewing set gives you roughly 32 scoreable proximal surfaces in a full adult dentition. In a stable NHS recall population, mostly adults seen every six to twelve months, with fluoride toothpaste and a history of restored-and-now-quiet disease, the prevalence of new, previously unflagged proximal lesions is low. Call it 3%. That’s about one lesion per patient, which matches what most principals would recognise from a morning of check-ups: you’re not finding new caries on everyone.
Per patient: 32 proximal surfaces
True lesions present (3%): ~1
Sound surfaces: ~31
Sensitivity 95% -> flags 0.95 of the 1 real lesion = 0.95 true positives
Specificity 85% -> misses 15% of 31 sound surfaces = 4.65 false positives
Boxes drawn per patient: 5.6
Boxes that are real: 0.95
Positive predictive value = 0.95 / 5.6 = 17%
Seventeen percent. Five out of every six boxes the software draws on that patient are wrong. Not because the model is bad, but because you have applied a 15% false alarm rate to thirty-one healthy surfaces and a 95% hit rate to one sick one. Thirty-one is a much bigger number than one, and that asymmetry eats the result alive.
Run the same model in the population it was validated on. Research datasets are enriched: investigators deliberately gather images with lesions, because a dataset of healthy teeth teaches a model nothing. Prevalence in those sets often sits at 20 to 30% of surfaces.
Same model, research dataset, 25% prevalence:
32 surfaces -> 8 lesions, 24 sound
True positives: 7.6
False positives: 3.6
PPV = 68%
Same software. Same threshold. Same sensitivity and specificity, unchanged. PPV goes from 17% to 68% purely because the disease moved. This is the entire argument, and it’s why the validation paper and your Tuesday morning can both be telling the truth while feeling like they describe different products.
What 17% does to a practice
The clinical consequence isn’t a wrong diagnosis. You’re a dentist, you look at the radiograph yourself, you dismiss the nonsense. The consequence is slower, and worse.
First, alarm fatigue sets in within about three weeks. When five of six boxes are noise, the brain stops treating boxes as information. Aviation and ITU alarm research has been saying this for decades and dentistry is not exempt. By week four the associate is clicking through the overlay without reading it, and at that point you’ve paid four figures a year for a tool that’s been mentally switched off. The one genuinely useful flag it produces gets dismissed alongside the rest.
Second, and more dangerous in an NHS or mixed setting: someone doesn’t dismiss them. The new associate, three months out of Sheffield, sees a red box on 26 distal with “84% confidence” next to it. The patient is in the chair, the software is the expensive thing the practice bought, and the box is right there on the screen the patient can also see. That’s a conversation with real gravitational pull toward intervention. In a UDA environment where diagnosis and treatment planning carry financial weight in both directions, and in a private conversion where a Band 2 becomes a composite, a 17% PPV overlay is an unhelpful thing to have hanging over the discussion. GDC cases have been built on less.
Third, patient-facing overlays are a distinct risk. Several of these tools market the colour-coded image as a case acceptance aid, and it works: patients accept more treatment when they see the AI agree with you. Fine, when the AI is right. Showing a patient five boxes and then talking four of them down is not a conversation that builds trust, and if you don’t talk them down, you have used a badly calibrated tool to sell dentistry.
Reading the threshold setting
Here’s the practical part. Most of these systems expose the threshold, though rarely on the front screen and rarely in plain language. What you’re looking for, in the admin or clinical settings panel:
| What the vendor calls it | What it actually is |
|---|---|
| Detection sensitivity: Low / Medium / High | The threshold. High = more boxes. |
| Confidence threshold (0.0–1.0) | The threshold, honestly labelled. |
| Display confidence minimum / Minimum confidence to display | Threshold applied at display time only. |
| Screening mode vs Diagnostic mode | Screening = low threshold. Diagnostic = higher. |
| Enamel lesion display: on / off | Not a threshold, but the single biggest driver of box count. |
That last row deserves its own paragraph. A large share of “false positives” in proximal caries AI are E1 and E2 lesions: real demineralisation confined to enamel, correctly identified, that you would monitor and not touch. The model isn’t wrong. It’s answering “is there demineralisation” when you’re asking “should I pick up a handpiece.” If your system lets you suppress enamel-only flags or display them in a separate, quieter colour, do that before you touch anything else. On many setups it halves the noise without losing a single dentine lesion.
Then ask your vendor these, in writing, before renewal:
What prevalence was the published sensitivity/specificity measured at? If they can’t say, or say “a large multi-site dataset,” press. The number you want is the proportion of positive surfaces in the test set. Anything above 15% means the headline PPV does not transfer to your recall list.
What is the default threshold, and is it the same one used in the validation study? These frequently differ. A model validated at 0.5 and shipped at 0.35 will behave meaningfully worse than its own paper.
Can I set different thresholds for different patient cohorts? High-caries-risk children in a deprived-area practice genuinely have 20%+ prevalence, and the high-sensitivity setting is correct for them. Your stable adult recalls do not. One practice, two thresholds, is the right answer and a few systems now support it.
Give me the confusion matrix, not the AUC. AUC of 0.92 tells you the model ranks well across all thresholds. It tells you nothing about the operating point you’ve actually been sold.
Recalibrating in your own practice
Don’t take anyone’s word for the numbers, including mine. Run a two-week audit, and make it cheap enough that it actually happens.
Pick 25 consecutive adult recall patients. For each bitewing set, have the examining dentist record their own diagnosis first, blind to the overlay, then reveal the AI. Log four counts: AI flagged and you agreed; AI flagged and you disagreed; AI silent and you found a lesion; AI silent and you agreed. That’s your confusion matrix, from your patients, on your default threshold.
Twenty-five patients at roughly 32 surfaces is 800 surfaces. Enough to see whether your false positive rate is 5% or 15%, which is the difference between a tool you’ll keep using and one that becomes wallpaper. Then raise the threshold one notch, run another 25, and compare. You are looking for the setting where you disagree with fewer than about one box in three. Most practices land somewhere near the middle setting, not the default, and the vendors know this perfectly well.
If you’re at the earlier stage of choosing between systems rather than tuning one you own, the broader evaluation criteria (regulatory class, PACS integration, how the outputs sit in your notes) are covered in our guide to AI radiograph reading.
One more thing worth saying plainly: raising the threshold means the software will miss lesions it would otherwise have caught. That is a real trade and you should make it consciously rather than by accident. The defence is that the AI is not the last line. You are. A tool at 80% sensitivity that you read every single time is worth more than a tool at 95% sensitivity that you stopped reading in March.
The vendors optimised for the number that sells. Your job on Monday is to work out which number keeps your associates looking at the screen.