Bitewings, Periapicals and Panoramics: Where AI Adds Most
Most practices switch on radiograph AI the way they’d switch on a new light: one button, everything illuminated at once. Bitewings, periapicals, panoramics, the odd occlusal, all piped through the same model, all coming back with boxes and confidence percentages. Six weeks later the nurse has muted the notifications, two associates have quietly stopped looking at the overlay, and the principal is wondering what the £300 a month is actually buying.
The fix is narrower than most vendors will tell you. Turn it on for bitewings. Leave it off, or at minimum leave it advisory-only, for everything else until you have local evidence it earns its place.
This isn’t vendor-agnostic hedging. It’s a claim about where the underlying detection problem is actually tractable, and the gap between view types is much wider than the marketing decks suggest.
Why bitewings are the easy case
A bitewing is a geometrically constrained image. The beam is roughly perpendicular to the contact points, the holder enforces a repeatable angle, the anatomy in frame is limited to crowns and crestal bone, and the pathology you care about most (interproximal caries) presents as a radiolucency in a predictable location. Thousands of bitewings taken in thousands of practices look broadly alike.
That consistency is what a convolutional model feeds on. When training data and clinical data share the same geometry, performance transfers. When they don’t, it degrades in ways that are hard to see from inside the surgery.
Compare that to a periapical. The operator is chasing an apex, the angulation shifts to clear a zygomatic buttress or a shallow palate, and the model now has to reason about root morphology, lamina dura continuity, periapical radiolucency versus normal anatomical structures (mental foramen, incisive canal, maxillary sinus recess), and post-treatment change. The variance in how the image was captured is enormous. Two PAs of the same LR6 taken twenty minutes apart by different operators can look like different teeth.
The panoramic is worse again. Ghost images, the cervical spine shadow, the palatoglossal air space, patient positioning errors that shift the focal trough, superimposition through the premolar region. A DPT is a reconstructed image of a curved surface flattened onto a plane. Asking a detection model to call an early interproximal lesion on one is asking it to work against the physics of how the image was made.
The numbers, as far as anyone publishes them
Published performance for AI caries detection on bitewings sits in a reasonably tight band. Across studies of the main commercial systems and academic models, sensitivity for proximal caries on bitewings typically lands between 0.80 and 0.92, with specificity around 0.85 to 0.95 and AUC values commonly reported at 0.90 and above. Overjet’s FDA clearance work and Pearl’s Second Opinion submissions both sit in that region. Studies of Overjet and Pearl versus dentist-only reading consistently show recall improvements in the 15 to 30 percent range for enamel-only and early dentinal lesions, which are exactly the lesions a busy clinician misses at the end of a list.
Move to periapical radiographs and the picture loosens. Periapical lesion detection studies report sensitivity more often in the 0.65 to 0.85 range, with much wider confidence intervals and far more disagreement between reference standards. Part of that is genuine model limitation. A large part is that the ground truth itself is contested: two experienced endodontists shown the same PA agree on the presence of apical periodontitis far less often than two dentists shown the same bitewings agree on an interproximal lesion. When your reference standard is noisy, your measured sensitivity is noisy too, and your real-world performance is anyone’s guess.
Panoramics are the widest spread of all. For the tasks pans are genuinely good at (tooth numbering, detecting impacted third molars, flagging retained roots, identifying gross bone loss patterns) models perform well, often above 0.90 for tooth detection and numbering. For caries on a pan, reported sensitivity drops into the 0.60s and 0.70s in several studies. That is not a screening tool. That is a coin-flip with a confidence score attached.
Here’s roughly how it stacks up, using mid-range published figures rather than best-case vendor numbers:
| View | Task | Typical sensitivity | Typical specificity | Verdict |
|---|---|---|---|---|
| Bitewing | Proximal caries | 0.80–0.92 | 0.85–0.95 | Switch on |
| Bitewing | Crestal bone level | 0.82–0.90 | 0.88–0.94 | Switch on |
| Periapical | Apical radiolucency | 0.65–0.85 | 0.75–0.90 | Advisory only |
| Periapical | Caries | 0.70–0.85 | 0.80–0.92 | Advisory only |
| Panoramic | Tooth numbering / impactions | 0.90+ | 0.90+ | Useful, narrow |
| Panoramic | Caries | 0.60–0.78 | 0.70–0.88 | Leave off |
The bitewing rows are the only ones where the numbers are consistent enough across studies, vendors and populations that you can plan a workflow around them.
A worked example from an eight-surgery mixed practice
Take a practice doing roughly 1,100 bitewing pairs a year across NHS and private lists. Assume a baseline of 4.2 interproximal lesions detected per 100 surfaces examined, which is in the normal range for an adult recall population with average caries risk.
Switch on AI at a 25 percent uplift in early-lesion recall, and you find an additional 1.05 lesions per 100 surfaces. Most of those are enamel-only or just into dentine. Clinically that is a prevention conversation and a fluoride varnish, not a restoration, and on an NHS band it changes your treatment plan not at all. What it changes is the conversation, the record, and the recall interval. Over a year across 1,100 bitewing pairs, you’re looking at roughly 80 to 120 additional lesions identified early enough to arrest.
Now run the same maths on periapicals. The same practice takes maybe 700 PAs a year, most of them diagnostic (a patient with pain, a pre-endo assessment, a post-op check). At a sensitivity of 0.72 and specificity of 0.82 against a low base rate of true apical pathology in that mixed population, you generate a substantial number of false positives. Say 15 percent of the PAs come back flagged for apical change that isn’t there. That’s over 100 flags a year, each of which costs an associate thirty seconds of looking, a moment of doubt, and occasionally a wholly unnecessary conversation with a patient about “something on the X-ray.”
That’s the actual cost of switching AI on everywhere. Not money. Attention, and the slow erosion of trust in the tool. An associate who has dismissed forty false apical flags stops reading the true one.
What this looks like on Monday morning
Most of the mainstream systems let you scope this. In Pearl Second Opinion, detection categories are configurable per image type in the admin panel. Overjet’s dashboard allows you to enable or suppress findings by category, and Dentiscope, VideaHealth and Diagnocat all offer some version of per-view or per-finding control. Check yours before you assume it’s all-or-nothing; the setting usually exists, it’s just not in the onboarding deck.
A sensible starting configuration:
BITEWING
proximal caries ON (overlay visible chairside)
crestal bone level ON (overlay visible chairside)
existing restorations ON (charting assist)
calculus ON
PERIAPICAL
caries ON (advisory panel, not overlay)
apical radiolucency ON (advisory panel, not overlay)
root morphology OFF
PANORAMIC
tooth numbering ON
impacted teeth ON
retained roots ON
caries OFF
periapical lesions OFF
bone loss (gross) ON (advisory panel)
The distinction between “overlay” and “advisory panel” matters more than it sounds. An overlay puts a coloured box on the image the patient can see on the chairside monitor. An advisory panel is a list the clinician can choose to open. For anything where you don’t trust the sensitivity, keep it off the image the patient is looking at. You do not want to be explaining a false positive to a nervous patient in real time, and you certainly don’t want a box appearing over a healthy apex while you’re mid-conversation about a crown.
Give it eight weeks on bitewings only. Log every AI flag your clinicians disagree with, in a shared sheet, with the tooth and surface. At week eight you will have between 150 and 400 data points from a practice of your size and you will know your own false-positive rate, which is the only number that actually matters for your patient population, your sensor and your exposure settings. Vendor figures come from other people’s images.
The objection you’ll hear from the sales team
They’ll say full-library activation gives you a fuller clinical picture, and that suppressing findings on pans means missing things. Two responses.
First, sensitivity you can’t trust isn’t a fuller picture, it’s noise wearing the costume of a picture. A pan caries flag at 0.68 sensitivity means the model misses roughly a third of what’s there while also flagging things that aren’t. You aren’t safer for having it on; you’ve just moved your uncertainty into a different format.
Second, the real risk in a UK practice isn’t the lesion you miss on a pan. It’s the clinician who stops engaging with the tool entirely because it cried wolf on the mandibular anteriors one too many times. Adoption is the scarce resource. Spend it where the evidence is strongest, build the habit, then widen.
There’s a good broader treatment of how these systems fit into NHS and mixed-practice workflow in our guide to AI radiograph reading, including the indemnity and record-keeping side, which is worth reading before you sign anything.
What to ask a vendor before you buy
Ask for per-view performance data, not aggregate. If a vendor quotes you a single sensitivity figure across their whole product, that figure is dominated by whichever view type made up most of their validation set, and you have no way of knowing which. Ask specifically: what is your sensitivity for proximal caries on bitewings, and what is it for the same finding on a DPT? A vendor who can answer that cleanly is one worth talking to further.
Ask what the validation population looked like. Models trained predominantly on US private-practice images from photostimulable phosphor plates behave differently on a UK NHS list shot on a five-year-old direct sensor with a different exposure protocol. Ask whether they’ve validated on UK data at all.
Ask what happens to the flags. GDC expectations and your indemnity position both assume a clinician made the diagnostic decision. If the software writes findings into the notes automatically, you need to know exactly how that reads to a reviewer three years later, and whether “AI-flagged, clinician disagreed” is a state your record can even represent.
And ask about the pans specifically. If the answer involves a lot of enthusiasm and no numbers, you’ve learned what you needed.
One more thing worth knowing
The bitewing-first approach has a second-order benefit nobody mentions: it makes your AI spend defensible at partner or DSO level. A pilot scoped to one view type, one finding category and one eight-week window produces an actual number. Additional lesions identified, clinician agreement rate, time per radiograph. Try to evaluate a full-library deployment and you’ll be arguing about vibes six months in, because the signal from the bitewings gets buried under the noise from everything else.
Start with the view where the physics, the geometry and the evidence all point the same direction. The rest can wait until your clinicians are asking for it, which, if the bitewing pilot goes well, they will be.