How AI Reads a Dental Radiograph, Explained Without the Maths
You do not need to understand convolutional layers to buy software well. You do need to understand what the software is actually looking at, because that single piece of knowledge is the difference between a clinician who trusts an AI output appropriately and one who either rubber-stamps it or dismisses it.
Most vendor demos skip this. You get shown a bitewing with tidy coloured boxes over three carious lesions, a confidence score, and a slide about how the model was trained on 12 million images. None of that tells you why the same model will flag a cervical burnout as caries on a poorly angulated premolar bitewing and then miss an obvious distal lesion on the tooth next to it. The answer is in how the machine sees, and how it sees is genuinely simple.
Your radiograph is a spreadsheet
A standard Carestream or Dürr intraoral sensor gives you an image that is, to the software, a grid of numbers. A typical size 2 sensor produces something in the region of 1300 × 1600 pixels. Each pixel holds a single greyscale value, usually 0 to 255 after processing, where 0 is black (air, the film edge, a large radiolucency) and 255 is white (amalgam, a post, an under-exposed area).
That is the entire input. No tooth notation, no knowledge that a tooth has enamel and dentine, no idea which way is up. Just a number per pixel.
Here is a tiny slice of what a region around an enamel-dentine junction looks like as numbers, across a lesion:
column → 140 138 141 139 137 142 140
139 141 138 136 135 140 139
142 139 118 102 99 115 138 ← values drop here
140 137 104 81 78 108 140
141 140 121 106 103 119 141
139 142 140 138 140 139 140
A dentist looking at that region on a monitor sees a dark triangle under the contact point. The model sees a cluster of values in the 78 to 121 band sitting inside a field of values in the 135 to 142 band. That local drop, its shape, its gradient at the edges, its position relative to other value clusters, is the whole signal.
Pattern detection, from edges upward
Training a model means showing it tens of thousands of these grids alongside labels a dentist has drawn, and letting it discover which arrangements of numbers correlate with which label. What it learns, in practice, builds up in layers of increasing size.
The earliest thing it learns is edges. A place where values go from 140 to 80 over three pixels is an edge. Learning to find edges in any orientation is trivial and the model nails it within the first few thousand training images. Then it learns corners, curves, and small textures: the particular gradient of enamel meeting air, the abrupt cliff of a restoration margin, the speckled texture of trabecular bone.
Further in, those small features get combined into bigger ones. A curve of a certain radius with bright values inside and mid values outside, repeated, becomes a recognisable crown outline. Two of those with a dark band between them becomes an interproximal contact. And a specific darkening shape, in a specific location relative to that contact, with a specific gradient at its borders, becomes what the label said was caries.
The model never learns the word “enamel.” It learns that a certain statistical arrangement of pixel values, appearing in a certain positional relationship to other arrangements, was labelled D2 caries 4,000 times in training. That is the whole mental model, and it explains almost every failure mode you will encounter.
What this predicts about failures, and why it matters at the chairside
Once you hold the pixel-pattern picture in your head, the trust question becomes tractable. Ask one thing of any AI output: is the pattern the model is responding to genuinely pathology, or is it a pixel arrangement that merely resembles one?
Cervical burnout is the cleanest example. On a bitewing where the beam has passed through a thinner section of the cervical region, you get a darkened band just below the CEJ. The pixel values drop. The gradient is soft rather than sharp, and the shape is a band rather than a wedge, but it is unquestionably a region of lower values in the place caries lives. Models flag it. Overjet Dental Assist, Pearl Second Opinion and VideaHealth have all published or acknowledged work on this specific class of error, and all three have improved on it, but none have eliminated it. When you see a flag at the cervical margin on a radiograph you know was angulated awkwardly, your prior should shift hard toward artefact.
Restoration overhangs produce the mirror problem. A bright amalgam edge creates an extreme value cliff, and immediately adjacent to a very bright region, sensors and processing pipelines produce a dark halo. That halo is a pixel pattern nearly identical to recurrent caries under a margin. This is why AI flags under existing restorations deserve more scepticism than flags on virgin surfaces, not less, which is the opposite of how most people instinctively read them.
The positive side is equally predictable. AI finds early interproximal lesions well because they are exactly the kind of small, low-contrast, locally-consistent value drop that a human eye on a busy Tuesday afternoon skims past. The published numbers here are real and they are not small. A 2021 study of Pearl’s system in Scientific Reports reported sensitivity gains of around 15 to 20 percentage points for caries detection over unaided dentists. VideaHealth’s FDA 510(k) submission data for VideaAI put caries detection sensitivity at 0.87 against a consensus ground truth where the mean individual dentist scored notably lower. Overjet has published internal figures showing roughly a 30% increase in detected proximal caries across participating practices. Treat any single vendor number as marketing, but the direction is consistent across independent work: machines are better than tired humans at spotting small dark patches, and worse at knowing whether a dark patch means anything.
The confidence score is not a probability of disease
This is where most purchasing conversations go wrong, so it is worth being blunt. When Pearl Second Opinion shows you 82%, that number is not “there is an 82% chance this tooth has caries.” It is closer to “82% of the way toward the pattern threshold this model learned from its labelled training set.”
Two things follow. First, the number is calibrated against the labellers, not against histology. If the training labels were drawn by a panel of three dentists reading radiographs, the model’s ceiling is agreement with three dentists reading radiographs, and inter-examiner agreement on D1 and D2 lesions is famously poor, with kappa values commonly landing between 0.4 and 0.6. The model cannot be more right than its teachers.
Second, the threshold is a business decision someone made. Most systems ship with a default that trades sensitivity against specificity, and many let you move it. Moving it down surfaces more lesions and more artefacts. Moving it up gives you cleaner output and misses more. For an NHS practice where a false positive means a difficult conversation about a Band 2 course of treatment the patient did not expect, and a false negative means a lesion caught six months later, that trade-off is a clinical governance question you should be deciding deliberately rather than inheriting from a vendor’s default. Ask in the demo whether the threshold is adjustable and who is permitted to adjust it. If the answer is vague, that tells you something.
A worked example from an actual appointment
Lower right, two bitewings, a 34-year-old patient with moderate plaque control. The software returns:
LR6 distal caries confidence 0.91 D2
LR5 mesial caries confidence 0.63 D1
LR4 cervical caries confidence 0.58 D1
LR7 distal no finding
Reading this with the pixel model in mind takes about eight seconds. The LR6 distal flag at 0.91 is a strong, well-defined value drop in the classic position, and your own eye confirms a clear dentinal shadow. Act on it.
The LR5 mesial at 0.63 sits in a genuinely ambiguous band. Look at whether the contact is open enough to see, whether the drop has a sharp enamel-side border, and whether the LR6 distal lesion’s own dark region is bleeding into the adjacent area on this projection. Monitor, fluoride varnish, re-image in six months.
The LR4 cervical flag at 0.58 is the one to interrogate hardest, because the pattern the model is reacting to is very likely geometric rather than pathological. Check the angulation. Check whether the darkening is a band following the root contour rather than a wedge. Nine times in ten this is burnout.
And the LR7 “no finding” is not a clearance. If that tooth is at the edge of the image, the model is working on partial data, and a distal lesion on the most posterior tooth in a bitewing is one of the more common misses in practice.
That reading process is not deep learning theory. It is one question, asked four times: what pixels is this responding to, and are those pixels disease?
What to actually ask a vendor
Bring five questions to the demo, and bring your own radiographs on a USB stick rather than looking at theirs.
| Question | What a good answer sounds like |
|---|---|
| Who drew the training labels, and how many per image? | Named number of clinicians, consensus process described, some histological validation |
| What was sensitivity and specificity, on whose data? | Separate figures, external validation set, not just internal test split |
| Is the confidence threshold adjustable? | Yes, with a described governance path |
| How does it behave on our sensor model? | They ask which sensor you run before answering |
| What happens to our images? | Clear UK GDPR position, named processor, DPIA support |
The sensor question catches more vendors out than you would expect. A model trained largely on Planmeca and Sirona output will behave measurably differently on a ten-year-old Schick sensor with a different noise profile, because the noise profile is part of the pixel pattern. Ask for a pilot on your own hardware, on your own images, for at least 200 radiographs before you commit to a contract. If a vendor will not do that, the problem is not your due diligence.
For the wider picture on selecting and governing these systems inside a UK practice, including the CQC and indemnity angles, our AI radiograph reading pillar goes considerably deeper than one post can.
Where the mental model stops being enough
Pixel-pattern thinking gets you a long way, and it is the right first tool. It will not tell you whether a model is drifting six months into deployment, whether your associates are gradually deferring to the software in a way that erodes their own reading, or whether a run of unexpected flags on one surgery’s images reflects real disease or a sensor that needs replacing. Those need audit, not intuition: pull a sample of 50 AI-flagged lesions per quarter, compare against what was actually found on opening, and write the number down.
The dentists who get the most out of these tools are not the ones who understand the architecture. They are the ones who stopped treating the output as an opinion from a colleague and started treating it as a measurement from an instrument, with a known range, known failure conditions, and a calibration record you keep yourself.