Human doctors have spent the past three years watching a tool arrive in their working lives faster than any technology before it. In January 2023, 38% of physicians surveyed by the American Medical Association used artificial intelligence professionally. By the AMA’s 2026 survey, fielded between 15 January and 2 February 2026 across 1,692 physicians, that figure was 81%.

Adoption curves like that are usually a story about convenience. This one is not. Alongside the adoption numbers sits a quieter finding: 88% of the same physicians reported concern about skill loss, particularly among early-career colleagues. The tool is being welcomed and mistrusted by the same people at the same time.

The question in the headline is not rhetorical hand-wringing. Human doctors are asking it in specialty forums, in medical school admissions offices, and in the pages of medical journals — and it deserves a better answer than either “AI will replace you” or “nothing will change.”

This article works through what the evidence actually shows. It separates the benchmark results from the clinical results, sets out which tasks are genuinely compressing and which are not, examines the one published study that found regular AI use measurably degrading a clinician’s unassisted performance, and looks at what UK regulators have already decided about who carries the responsibility.

Some of what follows is uncomfortable for the reassurance industry that has grown up around medical AI. Some of it is uncomfortable for the disruption industry too. The honest position sits between them, it is narrower than either camp admits, and it leaves human doctors with a smaller but far more defensible territory than the one they held in 2023.

Why Human Doctors Are Asking the Question Now

human doctors vs ai what is left for us b solid keystone block

The trigger was not a single product launch. It was the speed at which several separate things became true at once.

Adoption moved from experiment to routine

The AMA’s figures show the average number of AI use cases per physician rising from 1.1 to 2.3 between 2023 and 2026. The most common single use — summarising medical research and standards of care — went from 13% of physicians in 2024 to 39% in 2026. Creating discharge instructions, care plans or progress notes reached 30%, and documenting billing codes, charts or visit notes reached 28%.

Doximity’s 2026 State of AI in Medicine report, which surveyed 3,151 physicians across 15 specialties, found adoption rising from 47% in its early-2025 wave to 63% in the wave run between November 2025 and January 2026. Among the human doctors who had adopted the technology, 88% reported using it daily and half reported using it several times a day.

Professional AI use among physicians, AMA survey
2023 38%
2026 81%
Doximity, early 2025 47%
Doximity, late 2025 to early 2026 63%

The benchmarks started beating clinicians in public

Diagnostic benchmark results moved from journal appendices to newspaper front pages during 2025 and 2026. The most quoted of them put a language model four times ahead of practising clinicians, and it circulated without the caveats attached.

The patients arrived first

The third pressure is not institutional at all. Patients are already consulting a chatbot before, instead of, and after the appointment — and human doctors are frequently the last to hear about it.

The training pipeline noticed

The AMA found that 92% of physicians want more education and training on AI, while 27% had received none from any source. A profession in which human doctors train for a decade before practising independently is unusually sensitive to a change nobody has taught it to handle.

What the Benchmarks Actually Showed Human Doctors

human doctors vs ai what is left for us c solid tuning fork

In June 2025, Microsoft AI published results for a system it called the MAI Diagnostic Orchestrator, or MAI-DxO. The number that travelled was 85.5% against 20%.

The 85.5% result

MAI-DxO was tested on 304 recent case records from the New England Journal of Medicine, restructured into a sequential diagnostic benchmark: the system starts with limited information, asks questions, orders tests, and works towards a diagnosis. Paired with OpenAI’s o3 model, it reached 85.5% accuracy. Configured for cost, it reached 80% while spending roughly 20% less on testing than physicians did.

The 20% those physicians scored

The comparison group was 21 practising physicians from the United States and the United Kingdom, each with five to twenty years of clinical experience. Their mean accuracy across completed cases was 20%. Microsoft is explicit about the conditions: they worked without access to colleagues, textbooks, or generative AI, a restriction the company described as enabling “a fair comparison to raw human performance.”

That framing is worth pausing on. Raw human performance is not what any patient receives. Human doctors practise inside a system of colleagues, references, second opinions, and follow-up appointments, and removing all of it produces a number that describes a laboratory condition rather than a clinic.

Why a NEJM case series is not a surgery

The case mix compounds the problem. NEJM clinicopathological conferences are selected precisely because they are difficult and instructive. They contain no healthy patients, no self-limiting illness, and no undifferentiated presentations. Microsoft itself notes the research is a demonstration, not approved for clinical use, and that testing on common everyday presentations is still required.

Diagnostic accuracy on 304 NEJM sequential cases
MAI-DxO with o3, accuracy setting 85.5%
MAI-DxO with o3, cost setting 80%
21 physicians, no colleagues or references 20%

The precedent nobody wants to repeat

There is form here. In 2016 Geoffrey Hinton told a Toronto seminar that “people should stop training radiologists now,” because within five years deep learning would do the job better. A decade later, diagnostic radiology residency programmes filled at roughly 97% in the 2025 Match, around half of US radiologist vacancies went unfilled in 2023, and imaging volumes continue to outrun the workforce. The prediction was not wrong about the technology. It was wrong about the job, and about how much of that job human doctors were actually doing when they read a scan.

What the headline saidWhat the study measured
AI is four times better than doctorsOne orchestrator scored 85.5% where 21 isolated clinicians scored 20%
Doctors were given a fair testThey were barred from colleagues, textbooks and search tools
These were ordinary casesThey were teaching cases chosen for difficulty and rarity
The system is ready for clinicsThe vendor calls it a demonstration, not approved for clinical use
Diagnosis is the whole jobDiagnosis is one step in a process the benchmark never enters

The Work Human Doctors Still Own Outright

human doctors vs ai what is left for us d solid stepped ladder

Strip out the benchmark theatre and a shorter, firmer list remains. These are not sentimental categories. They are the parts of clinical work that no current system performs, because they are not text-in, text-out problems.

Getting the history that is not in the record

A large language model can only reason over what it is given. The clinical history is not a document waiting to be retrieved; human doctors produce it, during the consultation, by deciding which question to ask next and noticing that the answer does not fit. The STAT analysis of consumer health AI captured this precisely: models named the correct diagnosis more than 90% of the time when handed a complete case, and failed to produce a comprehensive differential more than 80% of the time when given only what was available at the first visit.

Examining the undifferentiated patient

Most of medicine does not present as a case record. It presents as a person who feels unwell. Deciding whether that is anxiety, anaemia or an early malignancy usually begins with physical examination and pattern recognition across a body, not a transcript. Nothing currently deployed does this, which is why the first ten minutes of most appointments remain the exclusive territory of human doctors.

Deciding not to test

Diagnostic systems optimise for reaching an answer. A large part of the clinical skill human doctors develop is knowing when not to pursue one — when a further scan will produce an incidental finding, a cascade of follow-ups, and net harm. That judgement is a decision about a specific person’s life, and it is the opposite of what a benchmark rewards.

Carrying a decision that cannot be deferred

Someone has to act while the picture is incomplete. Human doctors do this constantly, and the discomfort of it is the job. A system that produces a ranked differential has not made a decision; it has produced an input to one.

The Accountability Gap Only Human Doctors Can Fill

human doctors vs ai what is left for us e solid hourglass

The most durable answer to “what’s left for us” is not a clinical skill at all. It is legal and ethical standing.

A computer cannot be held accountable

An internal IBM training document from 1979 put it in a single line: “A computer can never be held accountable, therefore a computer must never make a management decision.” Forty-seven years later, no jurisdiction has found a way around this in medicine. The obligation to the patient — fiduciary, professional, and in the UK enforceable by the regulator — attaches to a person.

Somebody signs the note

This is why the regulatory guidance discussed below keeps landing in the same place. Whatever drafted the record, the clinician who signs it owns it. That single fact converts every AI tool in the consultation from a decision-maker into an instrument, regardless of how well it performs, and it puts human doctors at the end of every chain rather than beside it.

The liability question is still open

What human doctors consistently ask for is clarity on where legal liability sits when a tool they were told to use contributes to a bad outcome. The AMA found this among the top governance concerns, alongside the 55% of physicians who want direct involvement in AI implementation decisions at their own organisations. Adoption has moved considerably faster than that clarity has, and until it catches up the residual risk sits with human doctors by default.

The Deskilling Risk Human Doctors Should Take Seriously

human doctors vs ai what is left for us f solid drafting compass

The strongest argument that something real is being lost does not come from a think piece. It comes from an observational study in The Lancet Gastroenterology & Hepatology.

What the Polish colonoscopy study found

Researchers examined 1,443 unassisted colonoscopies performed at four centres in Poland taking part in the ACCEPT trial, before and after those centres introduced computer-aided polyp detection — a computer vision system that highlights suspicious tissue on the screen in real time. The endoscopists were experienced human doctors, each having performed more than 2,000 procedures. Nineteen took part.

Adenoma detection rate on standard, non-AI colonoscopy fell from 28.4% before the technology was introduced to 22.4% afterwards: a 6.0 percentage-point absolute drop, roughly a 20% relative reduction. AI-assisted procedures in the same period sat at 25.3%.

Adenoma detection rate before and after AI exposure, four Polish centres
Unassisted, before AI introduced 28.4%
AI-assisted procedures 25.3%
Unassisted, after AI exposure 22.4%

How the authors themselves framed it

Dr Marcin RomaÅ„czyk of the Academy of Silesia said the work was, to the team’s knowledge, “the first study to suggest a negative impact of regular AI use on healthcare professionals’ ability to complete a patient-relevant task in medicine of any kind.” Dr Omer Ahmad of University College London added that the findings “temper the current enthusiasm for rapid adoption of AI based technologies.”

The caveats that matter

This was observational, not randomised. Confounding factors unrelated to the technology cannot be excluded, and the finding applies to highly experienced operators — the effect on trainees who never practised unassisted is unknown, and is arguably the more important question.

The profession already suspected it

The AMA’s 88% skill-loss figure was collected independently, and it points at the same worry. When human doctors say they are concerned about early-career colleagues, this study is the shape of what they mean.

What AI Is Genuinely Taking From Human Doctors

Not everything in the disruption story is wrong. Some work is being removed, and it is worth being precise about which.

Documentation, partially

Ambient scribes are the real deployment story, and the clearest case of software taking work away from human doctors rather than competing with them. A study of roughly 1,800 clinicians across five academic medical centres between 2023 and 2025 found users saved 16 minutes of documentation time and spent 13 fewer minutes in the medical record for every eight hours of patient care. A UCLA randomised trial of one product measured a reduction of about 41 seconds per note.

Those are modest numbers against the promise. They also sit alongside a multicentre finding of a 31% reduction in reported burnout among scribe users, and Doximity’s finding that current users estimate a 48% cut in weekly after-hours charting. Time saved and strain relieved are not the same measurement, and the gap between them is where most of the marketing lives.

Literature search and summarisation

This is the largest single use case in both surveys, and it is a genuine substitution. The work of finding and condensing the current standard of care is now materially faster.

The routine, protocol-driven tail

Highly standardised outpatient work, protocol-driven telemedicine and repetitive administrative tasks are the parts most credibly compressing. That is a real change to the shape of some careers for human doctors, and it is not the same as replacing the profession.

TaskStatus in 2026Who carries it
Drafting the consultation noteLargely automatable todayTool drafts, clinician signs
Summarising the evidence baseSubstituted at scaleTool, with clinician verification
Generating a differential from a full caseStrong in benchmarksShared
Eliciting the history in the first placeNo deployed system does thisHuman doctors
Physical examinationNot addressedHuman doctors
Deciding to stop investigatingActively counter-incentivisedHuman doctors
Accountability for the outcomeLegally unassignable to softwareHuman doctors

What Regulators Have Settled for Human Doctors So Far

If the deployment is this fast, the evidence base underneath it deserves a look.

What the FDA list actually contains

As of 30 March 2026 the US Food and Drug Administration’s list of AI-enabled medical devices held 1,524 entries. Radiology accounts for 1,164 of them — around 76% of everything authorised. GE HealthCare leads by count with 130 authorisations, followed by Siemens Healthineers at 95 and Philips at 58.

The evidence behind the clearances

A cross-sectional review of 691 FDA-cleared devices found that only 1.6% cited data from randomised clinical trials, and fewer than 1% reported actual patient health outcomes. Clearance is a regulatory status, not a claim that the tool improves how patients do.

What the UK has decided about scribes

On 29 July 2026 the MHRA published guidance confirming that ambient voice technology products used solely for transcription, summarising a clinical conversation, drafting correspondence, or suggesting clinical codes for review are not medical devices under current UK rules. NHS England has issued companion guidance on safe adoption, with more due through 2026 and 2027.

The rollout underneath the guidance

The South West London Acute Provider Collaborative announced a contract with Lyrebird Health on 7 April to deploy ambient scribing to 10,000 clinicians in the first year, scaling towards 20,000 over four years. In both the UK and Australia, responsibility for the accuracy of the clinical record remains with the clinician, whatever drafted it.

The Patients Are Already Going Around Human Doctors

The pressure that most changes the answer to “what’s left” is not coming from hospital procurement. It is coming from outside the system entirely.

Forty million questions a day

By STAT’s account, more than 40 million Americans ask ChatGPT health questions every day, most of it outside clinic hours and beyond the sight of human doctors, and the medical disclaimers that once wrapped those answers have largely disappeared. Consumer platforms have grown alongside it: Function Health was valued at $2.5 billion in November 2026, and Doctronic reports 24 million consultations.

The shadow system and its blind spot

Arya Rao of Harvard Medical School and MIT and Dr Marc Succi of Mass General Brigham describe this as a shadow medical system: patients assembling diagnostics, prescriptions and consultations without a clinician in the loop. Their central objection is not accuracy. It is that a medical decision is made “almost by definition” in a moment of uncertainty and anxiety, and the entity answering carries no obligation to the person asking.

What it does inside the consultation

The AMA data shows the profession is barely sighted on this. Some 29% of physicians said no patient had ever disclosed using an AI chatbot, while 30% believed at least half of their patients were doing so. Notably, 70% supported patients using general AI tools for routine health information — the objection is to the invisibility, not the use.

The role this creates rather than removes

An unreviewed answer that arrived at 2am is now a routine input to the consultation. Interpreting it, correcting it and deciding what to do about it is new clinical work, and it lands on human doctors by default.

What Human Doctors Should Actually Do About It

None of this resolves into reassurance. It resolves into a short list of things worth doing while the picture is still forming.

Take the governance seat

Only 55% of physicians report wanting involvement in AI implementation decisions, and fewer still get it. Tools chosen without input from human doctors become tools that reshape clinical work without anyone deciding that they should. Organisations building an AI strategy around clinical services need practising clinicians in the procurement room, not in the feedback survey afterwards.

Fix the training gap deliberately

That 92% of physicians want more AI training while 27% have had none is the most actionable statistic in any of these surveys. It is also the cheapest to address, and it sits squarely in the ordinary discipline of artificial intelligence and machine learning deployment: nobody ships a system into daily use without training the people who depend on it.

Keep an unassisted baseline

The colonoscopy finding suggests a practical countermeasure that costs nothing: human doctors should retain regular unassisted practice, and measure it. If the skill is decaying, that shows up in the data long before it shows up in an outcome.

Insist on the outcome question

Given that fewer than 1% of cleared devices report patient outcomes, “does this change what happens to patients?” is a fair and currently unanswered procurement question, and human doctors are the people best placed to keep asking it. Our AI models, tools and releases hub tracks how vendors describe these systems as they ship.

Treat the record as a governed asset

Ambient capture puts a live recording of a clinical conversation through a third-party system. That is a data protection question as much as a clinical one, and it is easier to settle before deployment than after.

Human Doctors and AI: Frequently Asked Questions

Will AI replace human doctors?

Nothing in the current evidence supports it. Diagnostic benchmarks measure one step of clinical work under conditions no clinic reproduces, and the accountability that defines the role cannot presently be assigned to software. The credible change is to the composition of the job human doctors do, not to its existence.

Is the 85.5% versus 20% comparison fair?

It is accurately reported and badly framed. The physicians involved were denied colleagues, textbooks and search tools to isolate unaided human performance, and the cases were selected for difficulty. Both facts inflate the gap relative to real practice.

Does using AI actually make human doctors worse?

One observational study of 19 experienced endoscopists found unassisted adenoma detection fell 6.0 percentage points after AI was introduced. That is a real signal from a single non-randomised study, and it has not yet been replicated. It is a reason for monitoring, not for refusal.

Which specialties are changing fastest?

Radiology dominates the regulatory picture, with 1,164 of 1,524 FDA-listed AI-enabled devices. On the adoption side Doximity found neurologists highest at 64%, then gastroenterologists at 61% and internists at 60%.

How much time do AI scribes really save?

Roughly 16 minutes of documentation per eight hours of patient care in the largest multi-centre study to date, with one randomised trial measuring about 41 seconds per note. Reported burnout improvements are considerably larger than the time savings, which is worth understanding before setting expectations.

What should human doctors ask a vendor?

Whether the tool has been evaluated against patient outcomes rather than task accuracy, who is liable if it contributes to harm, what happens to the recording, and what unassisted performance is expected to look like after a year of use.

References and Further Reading