GDC Standards and AI-Assisted Diagnosis: Where Responsibility Sits
Search the GDC’s Standards for the Dental Team for the word “algorithm” and you get nothing. No mention of artificial intelligence, machine learning, or software. Some principals read that silence as a gap, something the regulator hasn’t got round to yet. It isn’t a gap. The standards are written to be indifferent to what tools you use, and that indifference is what decides the question of gdc standards ai diagnosis before the conversation even starts: the registrant who signs the notes owns the clinical decision, and the software owns nothing at all.
Standard 7.2 says work within your knowledge, skills, professional competence and abilities. Standard 7.1 says provide good quality care based on current evidence and authoritative guidance. Neither of those obligations transfers to a vendor when you tick “accept” on an overlay. If a case goes to a Practice Committee, the registrant in the room will be you, and the question will be what a reasonable dentist would have done with the radiograph in front of them. “Pearl flagged it” is not a defence. Neither, and this is the part people miss, is “Pearl didn’t flag it.”
The arithmetic makes disagreement your normal state
Run some real numbers on a typical mixed practice. Say you see 60 recall patients a week and take four bitewings at each, so 240 radiographs. Each bitewing gives you roughly ten scoreable proximal surfaces, which is 2,400 surfaces a week going through the model.
Take a caries detection tool performing at 90% sensitivity and 85% specificity, which is in the region the published work on Pearl Second Opinion, Overjet and VideaHealth sits for proximal caries on bitewings. Now assume 4% of those surfaces have radiographically detectable caries, reasonable for a stable low-risk adult recall list. That’s 96 diseased surfaces.
Diseased surfaces: 96 → 90% sensitivity → 86 true positives
Sound surfaces: 2,304 → 15% false pos. → 346 false positives
Total flags: 432
Positive predictive value: 86 / 432 = 20%
Four out of every five boxes drawn on your screen are not disease you would restore. That’s not a broken product. That’s what any detector does when it hunts a low-prevalence finding across a large number of opportunities, and the vendors know it. It means you will silently overrule the software several hundred times a week, and any documentation policy that treats each of those as a recordable event will collapse inside a fortnight.
So the governance question is not “how do I record disagreement.” It’s “which disagreements are material?”
Three tiers, and only two of them touch the notes
Tier one: routine non-agreement. The tool boxes an enamel shadow at the distal of UR6. You look, the marginal ridge is intact, the patient is caries-free for six years, fluoride varnish and a 12-month recall. You do not chart this. Charting it 346 times a week degrades your records, and Standard 4.1 asks for complete and accurate records, not exhaustive ones. What you do instead is set your detection threshold deliberately, at practice level, and write that decision down once.
Tier two: material disagreement. The software identifies something a reasonable colleague might have acted on, and you decide not to act. Periapical radiolucency at UL5, previously root treated, and you read it as a healing lesion rather than active pathology. Or bone loss quantified at 4.2 mm mesial of LR6 when your probing depths say 3 mm and the patient has a shallow film angle. This goes in the notes, with your reasoning, on the day.
Tier three: the tool was wrong in the dangerous direction. It saw nothing and you saw something. This goes in the notes, and it goes somewhere else too, which I’ll come to.
Here is wording that holds up when read back three years later by someone who wasn’t there:
26/09/26 Bitewings R+L (IRMER justification: 12/12 recall, high-risk
historic, prev restorations UR/UL quadrants).
Second Opinion overlay reviewed with pt on screen.
AI flagged: UR6-D enamel caries (conf. displayed 0.71),
UR5-M enamel caries (0.64).
Clinical: UR6-D marginal ridge intact, no cavitation on
direct vision, no catch, no symptoms. UR5-M not reproduced
on repeat angulation, judged cervical burnout.
Decision: monitor both, no restorative intervention.
Duraphat 22600 applied. Reviewed OH, interdental brushing
demonstrated. Rv 6/12 with repeat BWs, pt aware to report
sensitivity sooner.
Explained to pt that software marks areas for my attention
and that I do not agree these need filling now. Pt content.
Four things that note does. It names the tool and what it actually output, so nobody later argues about what you were shown. It gives your clinical findings separately from the software’s, which is the whole point. It states a decision and a review interval, so monitoring reads as a plan rather than an omission. And it records that the patient was told, which matters more than most people expect.
When the patient has already seen the overlay
Showing patients the coloured boxes is one of the strongest selling points these products have, and case acceptance data from the vendors leans on it heavily. It also creates a consent problem that didn’t exist five years ago.
Standard 2.3 requires you to give patients information in a way they can understand so they can make informed decisions. A patient who has watched software draw a red box on their tooth and then hears “we’ll leave that one” has received two contradictory messages, and the machine’s version looked more authoritative. If you don’t close that loop in the room, you have handed a future complaint its opening paragraph: the computer found a cavity and my dentist ignored it.
Say it out loud, in plain terms, every time you overrule a visible flag. Then write that you said it. It takes eleven seconds and it is the single highest-value line in the note.
Triage, the front desk, and the same principle wearing different clothes
Radiograph AI is where the debate lives, but the sharper risk in most NHS and mixed practices is the software nobody calls clinical. Symptom checkers in patient-facing booking, automated urgency scoring in an online triage form, recall prioritisation inside Dentally or SOE Exact, ambient note tools like Kiroku writing up your clinical record from the conversation.
Consider a real failure mode. A patient submits an online form describing swelling and difficulty swallowing. The triage logic scores it “routine” because “swelling” without “fever” fails a rule, and offers the next available slot in nine days. Nobody with a registration number read it. When that patient ends up in A&E, the GDC’s interest lands on the registrant responsible for the system: Standard 8.1, always put patients’ safety first, and Standard 6.3 on delegating appropriately. A nurse or receptionist acting on an algorithm’s urgency score is being asked to make a clinical judgement they are not registered to make, which puts them in breach of 7.2 through no fault of their own.
The fix is dull and cheap. Any automated triage output that reduces urgency gets read by a clinician before the patient is booked. Increases in urgency can pass through automatically, because the failure mode there is a wasted slot rather than a spreading infection. Write the rule down, name who reviews, and audit it monthly against actual attendances.
The three things to have in place before the next check
A written threshold decision. Whatever sensitivity setting your radiograph tool defaults to, someone clinical should have chosen it, dated the choice, and said why. That document is your answer to “you were overriding it constantly.”
A reporting route for tier three. AI diagnostic software in the UK is a regulated medical device requiring UKCA or CE marking, and since 16 June 2025 manufacturers have sat under strengthened post-market surveillance obligations in Great Britain. A tool that misses pathology is a device incident, reportable to the MHRA through the Yellow Card scheme as well as to the vendor. Almost nobody does this. It takes four minutes, and it is the difference between a practice that spotted a pattern and a practice that absorbed thirty near misses in silence. The wider compliance picture, including DPIA obligations under UK GDPR and where CQC Regulation 17 bites, sits in our Regulation, Data and Governance pillar.
A call to your indemnifier. Dental Protection, MDDUS and the DDU each take a view on AI-assisted reading, and those views are not identical. Ask specifically whether your cover responds if a claim alleges you relied on software output, and whether it responds if the allegation is that you ignored it. Get the answer in writing and keep it with your £680 ARF paperwork.
One more thing worth sitting with. IR(ME)R 2017 puts justification of every exposure on the named practitioner, and no software on the market can hold that role. When a tool’s business case depends on radiographs it hasn’t justified and cannot justify, the person who has to say no is you, on a Tuesday afternoon, with the rep still in the building.