Human doctors have spent the past three years watching a tool arrive in their working lives faster than any technology before it. In January 2023, 38% of physicians surveyed by the American Medical Association used artificial intelligence professionally. By the AMA’s 2026 survey, fielded between 15 January and 2 February 2026 across 1,692 physicians, that figure was 81%.
Adoption curves like that are usually a story about convenience. This one is not. Alongside the adoption numbers sits a quieter finding: 88% of the same physicians reported concern about skill loss, particularly among early-career colleagues. The tool is being welcomed and mistrusted by the same people at the same time.
The question in the headline is not rhetorical hand-wringing. Human doctors are asking it in specialty forums, in medical school admissions offices, and in the pages of medical journals — and it deserves a better answer than either “AI will replace you” or “nothing will change.”
This article works through what the evidence actually shows. It separates the benchmark results from the clinical results, sets out which tasks are genuinely compressing and which are not, examines the one published study that found regular AI use measurably degrading a clinician’s unassisted performance, and looks at what UK regulators have already decided about who carries the responsibility.
Some of what follows is uncomfortable for the reassurance industry that has grown up around medical AI. Some of it is uncomfortable for the disruption industry too. The honest position sits between them, it is narrower than either camp admits, and it leaves human doctors with a smaller but far more defensible territory than the one they held in 2023.
Table of contents
- Why Human Doctors Are Asking the Question Now
- What the Benchmarks Actually Showed Human Doctors
- The Work Human Doctors Still Own Outright
- The Accountability Gap Only Human Doctors Can Fill
- The Deskilling Risk Human Doctors Should Take Seriously
- What AI Is Genuinely Taking From Human Doctors
- What Regulators Have Settled for Human Doctors So Far
- The Patients Are Already Going Around Human Doctors
- What Human Doctors Should Actually Do About It
- Human Doctors and AI: Frequently Asked Questions
- References and Further Reading
Why Human Doctors Are Asking the Question Now
The trigger was not a single product launch. It was the speed at which several separate things became true at once.
Adoption moved from experiment to routine
The AMA’s figures show the average number of AI use cases per physician rising from 1.1 to 2.3 between 2023 and 2026. The most common single use — summarising medical research and standards of care — went from 13% of physicians in 2024 to 39% in 2026. Creating discharge instructions, care plans or progress notes reached 30%, and documenting billing codes, charts or visit notes reached 28%.
Doximity’s 2026 State of AI in Medicine report, which surveyed 3,151 physicians across 15 specialties, found adoption rising from 47% in its early-2025 wave to 63% in the wave run between November 2025 and January 2026. Among the human doctors who had adopted the technology, 88% reported using it daily and half reported using it several times a day.
The benchmarks started beating clinicians in public
Diagnostic benchmark results moved from journal appendices to newspaper front pages during 2025 and 2026. The most quoted of them put a language model four times ahead of practising clinicians, and it circulated without the caveats attached.
The patients arrived first
The third pressure is not institutional at all. Patients are already consulting a chatbot before, instead of, and after the appointment — and human doctors are frequently the last to hear about it.
The training pipeline noticed
The AMA found that 92% of physicians want more education and training on AI, while 27% had received none from any source. A profession in which human doctors train for a decade before practising independently is unusually sensitive to a change nobody has taught it to handle.
What the Benchmarks Actually Showed Human Doctors
In June 2025, Microsoft AI published results for a system it called the MAI Diagnostic Orchestrator, or MAI-DxO. The number that travelled was 85.5% against 20%.
The 85.5% result
MAI-DxO was tested on 304 recent case records from the New England Journal of Medicine, restructured into a sequential diagnostic benchmark: the system starts with limited information, asks questions, orders tests, and works towards a diagnosis. Paired with OpenAI’s o3 model, it reached 85.5% accuracy. Configured for cost, it reached 80% while spending roughly 20% less on testing than physicians did.
The 20% those physicians scored
The comparison group was 21 practising physicians from the United States and the United Kingdom, each with five to twenty years of clinical experience. Their mean accuracy across completed cases was 20%. Microsoft is explicit about the conditions: they worked without access to colleagues, textbooks, or generative AI, a restriction the company described as enabling “a fair comparison to raw human performance.”
That framing is worth pausing on. Raw human performance is not what any patient receives. Human doctors practise inside a system of colleagues, references, second opinions, and follow-up appointments, and removing all of it produces a number that describes a laboratory condition rather than a clinic.
Why a NEJM case series is not a surgery
The case mix compounds the problem. NEJM clinicopathological conferences are selected precisely because they are difficult and instructive. They contain no healthy patients, no self-limiting illness, and no undifferentiated presentations. Microsoft itself notes the research is a demonstration, not approved for clinical use, and that testing on common everyday presentations is still required.
The precedent nobody wants to repeat
There is form here. In 2016 Geoffrey Hinton told a Toronto seminar that “people should stop training radiologists now,” because within five years deep learning would do the job better. A decade later, diagnostic radiology residency programmes filled at roughly 97% in the 2025 Match, around half of US radiologist vacancies went unfilled in 2023, and imaging volumes continue to outrun the workforce. The prediction was not wrong about the technology. It was wrong about the job, and about how much of that job human doctors were actually doing when they read a scan.
| What the headline said | What the study measured |
|---|---|
| AI is four times better than doctors | One orchestrator scored 85.5% where 21 isolated clinicians scored 20% |
| Doctors were given a fair test | They were barred from colleagues, textbooks and search tools |
| These were ordinary cases | They were teaching cases chosen for difficulty and rarity |
| The system is ready for clinics | The vendor calls it a demonstration, not approved for clinical use |
| Diagnosis is the whole job | Diagnosis is one step in a process the benchmark never enters |
The Work Human Doctors Still Own Outright
Strip out the benchmark theatre and a shorter, firmer list remains. These are not sentimental categories. They are the parts of clinical work that no current system performs, because they are not text-in, text-out problems.
Getting the history that is not in the record
A large language model can only reason over what it is given. The clinical history is not a document waiting to be retrieved; human doctors produce it, during the consultation, by deciding which question to ask next and noticing that the answer does not fit. The STAT analysis of consumer health AI captured this precisely: models named the correct diagnosis more than 90% of the time when handed a complete case, and failed to produce a comprehensive differential more than 80% of the time when given only what was available at the first visit.
Examining the undifferentiated patient
Most of medicine does not present as a case record. It presents as a person who feels unwell. Deciding whether that is anxiety, anaemia or an early malignancy usually begins with physical examination and pattern recognition across a body, not a transcript. Nothing currently deployed does this, which is why the first ten minutes of most appointments remain the exclusive territory of human doctors.
Deciding not to test
Diagnostic systems optimise for reaching an answer. A large part of the clinical skill human doctors develop is knowing when not to pursue one — when a further scan will produce an incidental finding, a cascade of follow-ups, and net harm. That judgement is a decision about a specific person’s life, and it is the opposite of what a benchmark rewards.
Carrying a decision that cannot be deferred
Someone has to act while the picture is incomplete. Human doctors do this constantly, and the discomfort of it is the job. A system that produces a ranked differential has not made a decision; it has produced an input to one.
The Accountability Gap Only Human Doctors Can Fill
The most durable answer to “what’s left for us” is not a clinical skill at all. It is legal and ethical standing.
A computer cannot be held accountable
An internal IBM training document from 1979 put it in a single line: “A computer can never be held accountable, therefore a computer must never make a management decision.” Forty-seven years later, no jurisdiction has found a way around this in medicine. The obligation to the patient — fiduciary, professional, and in the UK enforceable by the regulator — attaches to a person.
Somebody signs the note
This is why the regulatory guidance discussed below keeps landing in the same place. Whatever drafted the record, the clinician who signs it owns it. That single fact converts every AI tool in the consultation from a decision-maker into an instrument, regardless of how well it performs, and it puts human doctors at the end of every chain rather than beside it.
The liability question is still open
What human doctors consistently ask for is clarity on where legal liability sits when a tool they were told to use contributes to a bad outcome. The AMA found this among the top governance concerns, alongside the 55% of physicians who want direct involvement in AI implementation decisions at their own organisations. Adoption has moved considerably faster than that clarity has, and until it catches up the residual risk sits with human doctors by default.
The Deskilling Risk Human Doctors Should Take Seriously
The strongest argument that something real is being lost does not come from a think piece. It comes from an observational study in The Lancet Gastroenterology & Hepatology.
What the Polish colonoscopy study found
Researchers examined 1,443 unassisted colonoscopies performed at four centres in Poland taking part in the ACCEPT trial, before and after those centres introduced computer-aided polyp detection — a computer vision system that highlights suspicious tissue on the screen in real time. The endoscopists were experienced human doctors, each having performed more than 2,000 procedures. Nineteen took part.
Adenoma detection rate on standard, non-AI colonoscopy fell from 28.4% before the technology was introduced to 22.4% afterwards: a 6.0 percentage-point absolute drop, roughly a 20% relative reduction. AI-assisted procedures in the same period sat at 25.3%.
How the authors themselves framed it
Dr Marcin RomaÅ„czyk of the Academy of Silesia said the work was, to the team’s knowledge, “the first study to suggest a negative impact of regular AI use on healthcare professionals’ ability to complete a patient-relevant task in medicine of any kind.” Dr Omer Ahmad of University College London added that the findings “temper the current enthusiasm for rapid adoption of AI based technologies.”
The caveats that matter
This was observational, not randomised. Confounding factors unrelated to the technology cannot be excluded, and the finding applies to highly experienced operators — the effect on trainees who never practised unassisted is unknown, and is arguably the more important question.
The profession already suspected it
The AMA’s 88% skill-loss figure was collected independently, and it points at the same worry. When human doctors say they are concerned about early-career colleagues, this study is the shape of what they mean.
What AI Is Genuinely Taking From Human Doctors
Not everything in the disruption story is wrong. Some work is being removed, and it is worth being precise about which.
Documentation, partially
Ambient scribes are the real deployment story, and the clearest case of software taking work away from human doctors rather than competing with them. A study of roughly 1,800 clinicians across five academic medical centres between 2023 and 2025 found users saved 16 minutes of documentation time and spent 13 fewer minutes in the medical record for every eight hours of patient care. A UCLA randomised trial of one product measured a reduction of about 41 seconds per note.
Those are modest numbers against the promise. They also sit alongside a multicentre finding of a 31% reduction in reported burnout among scribe users, and Doximity’s finding that current users estimate a 48% cut in weekly after-hours charting. Time saved and strain relieved are not the same measurement, and the gap between them is where most of the marketing lives.
Literature search and summarisation
This is the largest single use case in both surveys, and it is a genuine substitution. The work of finding and condensing the current standard of care is now materially faster.
The routine, protocol-driven tail
Highly standardised outpatient work, protocol-driven telemedicine and repetitive administrative tasks are the parts most credibly compressing. That is a real change to the shape of some careers for human doctors, and it is not the same as replacing the profession.
| Task | Status in 2026 | Who carries it |
|---|---|---|
| Drafting the consultation note | Largely automatable today | Tool drafts, clinician signs |
| Summarising the evidence base | Substituted at scale | Tool, with clinician verification |
| Generating a differential from a full case | Strong in benchmarks | Shared |
| Eliciting the history in the first place | No deployed system does this | Human doctors |
| Physical examination | Not addressed | Human doctors |
| Deciding to stop investigating | Actively counter-incentivised | Human doctors |
| Accountability for the outcome | Legally unassignable to software | Human doctors |
What Regulators Have Settled for Human Doctors So Far
If the deployment is this fast, the evidence base underneath it deserves a look.
What the FDA list actually contains
As of 30 March 2026 the US Food and Drug Administration’s list of AI-enabled medical devices held 1,524 entries. Radiology accounts for 1,164 of them — around 76% of everything authorised. GE HealthCare leads by count with 130 authorisations, followed by Siemens Healthineers at 95 and Philips at 58.
The evidence behind the clearances
A cross-sectional review of 691 FDA-cleared devices found that only 1.6% cited data from randomised clinical trials, and fewer than 1% reported actual patient health outcomes. Clearance is a regulatory status, not a claim that the tool improves how patients do.
What the UK has decided about scribes
On 29 July 2026 the MHRA published guidance confirming that ambient voice technology products used solely for transcription, summarising a clinical conversation, drafting correspondence, or suggesting clinical codes for review are not medical devices under current UK rules. NHS England has issued companion guidance on safe adoption, with more due through 2026 and 2027.
The rollout underneath the guidance
The South West London Acute Provider Collaborative announced a contract with Lyrebird Health on 7 April to deploy ambient scribing to 10,000 clinicians in the first year, scaling towards 20,000 over four years. In both the UK and Australia, responsibility for the accuracy of the clinical record remains with the clinician, whatever drafted it.
The Patients Are Already Going Around Human Doctors
The pressure that most changes the answer to “what’s left” is not coming from hospital procurement. It is coming from outside the system entirely.
Forty million questions a day
By STAT’s account, more than 40 million Americans ask ChatGPT health questions every day, most of it outside clinic hours and beyond the sight of human doctors, and the medical disclaimers that once wrapped those answers have largely disappeared. Consumer platforms have grown alongside it: Function Health was valued at $2.5 billion in November 2026, and Doctronic reports 24 million consultations.
The shadow system and its blind spot
Arya Rao of Harvard Medical School and MIT and Dr Marc Succi of Mass General Brigham describe this as a shadow medical system: patients assembling diagnostics, prescriptions and consultations without a clinician in the loop. Their central objection is not accuracy. It is that a medical decision is made “almost by definition” in a moment of uncertainty and anxiety, and the entity answering carries no obligation to the person asking.
What it does inside the consultation
The AMA data shows the profession is barely sighted on this. Some 29% of physicians said no patient had ever disclosed using an AI chatbot, while 30% believed at least half of their patients were doing so. Notably, 70% supported patients using general AI tools for routine health information — the objection is to the invisibility, not the use.
The role this creates rather than removes
An unreviewed answer that arrived at 2am is now a routine input to the consultation. Interpreting it, correcting it and deciding what to do about it is new clinical work, and it lands on human doctors by default.
What Human Doctors Should Actually Do About It
None of this resolves into reassurance. It resolves into a short list of things worth doing while the picture is still forming.
Take the governance seat
Only 55% of physicians report wanting involvement in AI implementation decisions, and fewer still get it. Tools chosen without input from human doctors become tools that reshape clinical work without anyone deciding that they should. Organisations building an AI strategy around clinical services need practising clinicians in the procurement room, not in the feedback survey afterwards.
Fix the training gap deliberately
That 92% of physicians want more AI training while 27% have had none is the most actionable statistic in any of these surveys. It is also the cheapest to address, and it sits squarely in the ordinary discipline of artificial intelligence and machine learning deployment: nobody ships a system into daily use without training the people who depend on it.
Keep an unassisted baseline
The colonoscopy finding suggests a practical countermeasure that costs nothing: human doctors should retain regular unassisted practice, and measure it. If the skill is decaying, that shows up in the data long before it shows up in an outcome.
Insist on the outcome question
Given that fewer than 1% of cleared devices report patient outcomes, “does this change what happens to patients?” is a fair and currently unanswered procurement question, and human doctors are the people best placed to keep asking it. Our AI models, tools and releases hub tracks how vendors describe these systems as they ship.
Treat the record as a governed asset
Ambient capture puts a live recording of a clinical conversation through a third-party system. That is a data protection question as much as a clinical one, and it is easier to settle before deployment than after.
Human Doctors and AI: Frequently Asked Questions
Will AI replace human doctors?
Nothing in the current evidence supports it. Diagnostic benchmarks measure one step of clinical work under conditions no clinic reproduces, and the accountability that defines the role cannot presently be assigned to software. The credible change is to the composition of the job human doctors do, not to its existence.
Is the 85.5% versus 20% comparison fair?
It is accurately reported and badly framed. The physicians involved were denied colleagues, textbooks and search tools to isolate unaided human performance, and the cases were selected for difficulty. Both facts inflate the gap relative to real practice.
Does using AI actually make human doctors worse?
One observational study of 19 experienced endoscopists found unassisted adenoma detection fell 6.0 percentage points after AI was introduced. That is a real signal from a single non-randomised study, and it has not yet been replicated. It is a reason for monitoring, not for refusal.
Which specialties are changing fastest?
Radiology dominates the regulatory picture, with 1,164 of 1,524 FDA-listed AI-enabled devices. On the adoption side Doximity found neurologists highest at 64%, then gastroenterologists at 61% and internists at 60%.
How much time do AI scribes really save?
Roughly 16 minutes of documentation per eight hours of patient care in the largest multi-centre study to date, with one randomised trial measuring about 41 seconds per note. Reported burnout improvements are considerably larger than the time savings, which is worth understanding before setting expectations.
What should human doctors ask a vendor?
Whether the tool has been evaluated against patient outcomes rather than task accuracy, who is liable if it contributes to harm, what happens to the recording, and what unassisted performance is expected to look like after a year of use.
References and Further Reading
More than 80% of physicians use AI professionally: AMA survey
Physician Survey on Augmented Intelligence — American Medical Association
Doximity 2026 State of AI in Medicine Report
The Path to Medical Superintelligence — Microsoft AI
Sequential Diagnosis with Language Models — Microsoft Research
Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy
Routine AI assistance may lead to loss of skills in health professionals who perform colonoscopies
AI has created a shadow medical system — STAT
Medical AI will rob physicians of what little autonomy they still have — STAT
New UK Guidance Clarifies Medical Device Status of AI Scribes
Numbers from the FDA show radiology is maintaining its lead — The Imaging Wire
Large AI scribe study finds modest time savings, inconsistent use — STAT
The Godfather of AI Predicted I Wouldn’t Have a Job. He Was Wrong. — The New Republic
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.