AI Measurement of Periodontal Bone Loss: Useful or Noise?
Every demo goes the same way. The rep loads a periapical of an upper first molar, clicks once, and the software drops a coloured line from the CEJ to the bone crest and prints “4.2 mm, 38% root length, Stage III”. Someone in the room says “that’s clever”. Then everyone looks at the clinician, who shrugs, because they already knew that tooth was in trouble before the image finished rendering.
That shrug is the honest response, and it is also why most practices who buy ai periodontal bone loss detection get disappointed. They bought it as a diagnostic aid for the appointment in front of them. It is not very good at that, and it does not need to be. The value sits somewhere else entirely: in comparing today’s bone level against the one from 2022 using the same method, and in making sure the stage you write in the notes is the stage your associate would have written, and the stage the periodontist receives.
You already made the diagnosis before the AI finished loading
Consider what actually happens in a UK chair. BPE comes first. A code 4 sextant triggers full six-point pocket charting, and by the time you have recorded an 8 mm pocket distobuccal of the UL6 with bleeding on probing, the radiograph is confirmatory rather than revelatory. Bone loss visible on a periapical is late news anyway: radiographic change typically lags real mineral loss substantially, and radiographs have long been shown to underestimate intrasurgical bone loss by roughly a millimetre on average. The probe got there first.
So when a vendor quotes accuracy figures for single-image detection, ask what the comparator is. Krois and colleagues, publishing in Scientific Reports in 2019, reported a convolutional network detecting periodontal bone loss on panoramic segments at about 0.81 accuracy against a reference standard, with six dentists sitting at roughly 0.76. That is a real result and it is worth knowing. It is also a margin of five percentage points on a task you are doing alongside a probe, a pocket chart, a smoking history and a HbA1c. Nobody’s treatment plan changes on five points.
Where the money actually is: the same ruler, every time
Now change the question. Instead of “is there bone loss”, ask “has this site got worse since February 2024, and by how much”.
That question is brutally hard to answer by eye, and it is the question that determines whether you monitor, re-treat, or refer. Two periapicals taken two years apart by two different nurses, on two different holders, at slightly different vertical angles, are not directly comparable by inspection. You are comparing apparent distances in projections that differ. Human intra-examiner error on manual radiographic bone level measurement is commonly reported around half a millimetre, and inter-examiner error is worse. Half a millimetre is most of a year’s progression in a stable-ish Grade B patient.
Automated measurement does not remove the geometry problem. What it removes is the inconsistency of the person holding the ruler. Same landmark definition, same algorithm, same output units, every image, every time. A systematic bias applied identically to both images largely cancels when you subtract.
Worked example. A 44-year-old ex-smoker, UL6, mesial site, root length measured at 13.0 mm:
DATE CEJ-CREST % ROOT LEN Δ mm vs prior Δ mm since baseline
2022-01-18 2.6 mm 20% – –
2024-02-11 3.3 mm 25% +0.7 +0.7
2026-09-14 4.4 mm 34% +1.1 +1.8
Look at what that table gives you that a side-by-side visual comparison does not. Total loss of 1.8 mm over 4.7 years, running at roughly 0.38 mm/year. The indirect grade calculation, 34% bone loss divided by age 44, comes out at 0.77, which lands in Grade B. Direct evidence of progression sits just under the 2 mm in 5 years that would push it to Grade C. The software will not resolve that tension for you and should not pretend to. But you are now arguing about a number with a provenance, in front of a patient who can see the trend, rather than squinting at two films and saying “hmm, looks a bit worse”.
That conversation converts. It also documents. If that patient later complains that perio was missed, a dated numeric series beats “monitored, no significant change” in the notes by a distance.
Staging consistency is a practice problem, not a clinical knowledge problem
The 2017 World Workshop classification, as implemented by the BSP, stages on interproximal bone loss at the worst site:
| Stage | Bone loss at worst site | Where it sits |
|---|---|---|
| I | < 15% of root length | Coronal third, early |
| II | 15% to 33% | Coronal third |
| III | > 33% | Mid third and beyond |
| IV | > 33% plus complexity factors | Apical third, ≤ 20 teeth remaining, masticatory dysfunction |
Grading then runs bone loss percentage over age: under 0.5 is A, 0.5 to 1.0 is B, over 1.0 is C.
Here is the practical failure. Stage II and Stage III are separated by a single number, 33%, applied to the worst site in the mouth. In a three-surgery practice with two associates, a therapist and a locum covering Thursdays, that number is being estimated by eye, differently, all week. Published agreement between general dentists on staging has repeatedly come out mediocre, often in the moderate-kappa range rather than the strong agreement you would want from something that drives referral thresholds and NHS course-of-treatment decisions. The disagreement is not about competence. It is about eyeballing a percentage.
Automating that one measurement, consistently, across every clinician in the building, is worth more than another few points of detection sensitivity. It means the Stage III patient gets the Stage III pathway regardless of who they happened to be booked with, and your referrals to the perio specialist stop bouncing back with “please clarify staging”.
What the tools do, and what to make them prove
Overjet holds FDA clearance covering quantified bone level measurement and pushes per-tooth millimetre output directly into the chart. Pearl’s Second Opinion covers bone loss among a broader pathology set and has regulatory clearances across multiple markets. VideaHealth offers a perio module in the same territory. Diagnocat will produce per-tooth periodontal reporting from both 2D and CBCT. These are genuinely different products with different integration stories, and the broader question of how they slot into your radiograph workflow is covered in our pillar on AI radiograph reading.
Four things to demand in the demo, none of which vendors volunteer:
Bitewing geometry. If your practice takes horizontal bitewings as standard, the apex is off the image, root length cannot be established, and percentage-based staging is impossible. The software can give you a millimetre figure but not a stage. Practices that want staging out of AI usually need to move to vertical bitewings or routine periapicals of the worst sextants, which is a nursing and consent change, not a software change.
Longitudinal output. Ask directly: does it store measurements as structured data linked to tooth and site, and can it produce a delta against a named prior date? Several tools will happily annotate today’s image and retain nothing comparable. That product is the expensive version of a shrug.
Failure behaviour. Watch it run on a film with a poorly defined CEJ, heavy calculus, an overhanging amalgam and a bit of cone cut. What does it print? Silence and a confidence flag is a good answer. A confident 3.8 mm is not.
UK compliance. UKCA or CE marking under the transitional arrangements, a completed DTAC if you have any NHS trust or ICB exposure, a DPIA you can actually read, and a clear answer on where images are processed and whether they are used for model training. Get it in writing before the DSPT toolkit renewal, not after.
Setting your own noise floor
Before the first invoice, write down what counts as real. A sensible starting position for most mixed practices: changes under 0.5 mm at a single site are noise and trigger nothing. Between 0.5 mm and 1.0 mm, flag for review at the next recall with a repeat radiograph using matched technique. Over 1.0 mm at any site, or over 0.5 mm at three or more sites, goes to a full reassessment appointment.
Those thresholds are a policy decision, not a technical one, and they belong to the practice principal. Set them too tight and your therapist’s diary fills with reassessments generated by angulation artefacts. Set them too loose and you have bought a tool that agrees with you.
Costs are the easy part of the calculation. Sub-based dental AI in the UK is commonly quoted somewhere between roughly £150 and £400 per surgery per month depending on modules and site count, so a three-surgery practice is looking at something in the region of £6,000 to £14,000 a year. Against that, a Band 2 course of treatment and a handful of retained private perio maintenance patients per month is not a difficult sum. Against single-visit diagnosis alone, it is a terrible sum, because you were already doing that part correctly.
Pull five patients from your recall list next week whose last perio radiographs are two or more years old. Run both sets through whatever you are trialling and put the deltas in a spreadsheet. If the trend lines tell you something you did not already know, you have found your use case. If they do not, tell the rep exactly that.