You are no longer being compared to the surgeon down the hall. You are being compared to the patient's brother at home with a ChatGPT Pro account.
In August, a JAMA perspective co-authored by Zeke Emanuel and Vinod Khosla argued that autonomous AI will likely be ready for some clinical workflows by 2030, and that “having humans in the loop will likely worsen patient care.”[1] The AMA's CEO, John Whyte, disagreed publicly within days.[2] The argument is now on the record.
And every physician I know has started, without quite saying so, trying to prove they still matter in their own clinic.
I think we do. But we need to explain what, specifically, we add, and learn how to add it.
What does the data say?
In a 2024 trial, fifty physicians got an hour to review clinical vignettes. Half could use ChatGPT. The group with the chatbot scored 76 percent; the group without it, 74 percent. No significant difference.[3] The headline was that GPT-4 alone beat both groups. That 16-percentage-point advantage over the physicians using conventional resources came from an exploratory analysis of three runs of the model. It was not the comparison the trial was built to test.[3]
A 2025 follow-up found that AI assistance improved physicians' management reasoning, but detected no significant difference between assisted physicians and AI alone.[4] This April, in Science, a reasoning model outperformed two attending internists on diagnostic assessments of emergency department cases, all working from the same chart text.[5]
Do these results show physician irrelevance? Look at what the trial setup guaranteed before the clock started. The question was chosen. The case arrived framed as a clinical problem. The available information was fixed. There was no patient to examine and reexamine, no colleague to call, no additional history to find.
None of these comparisons tested what a physician might add by seeing the patient, gathering new information, or working with the AI rather than beside it.[3][4][5] The Science authors say as much: text-based performance is a narrow measure in a profession full of nontext information.[5]
That distinction is the whole argument.
What a physician adds
Context. More specifically, knowing which context matters.
Watch a good intern work up a new consult. They gather everything: birth history, every prior operation, the medication list read aloud from the bag at the bottom of the wife's purse. Then watch an attending surgeon talk to the same patient for five minutes and walk out with what she needs to make a decision. Part of that expertise is knowing what to ask, because she has seen this problem enough times to know which facts change the plan.
Ask a model about a diabetic foot ulcer and you get a competent summary: perfusion, offloading, debridement, infection control, glucose. Perfusion is one item of five. For the patient in front of me with a monophasic pedal signal and delayed capillary refill, perfusion changes the urgency and the sequence of everything else. The model answered the question it was asked. Anyone can type a prompt. Not everyone knows which question this foot requires.
I would bet the last five notes on that foot say “warm and perfused,” and at least one documents a pulse that does not exist. A model reading that chart inherits the error. You can feel the foot. You may hear something in how the patient describes the pain that changes the next question you ask.
Is context worth anything? Economists gave 227 radiologists AI assistance reading images. On average, it did not help. Providing clinical history and other patient information did improve their accuracy.[6] I suspect the value of being in the room goes further than a history field. This is what we were taught on day one of clinical rotations: just go see the patient.
The second skill
Context does not, by itself, optimize care. The same radiology experiment found that physicians under-trusted the AI and failed to account for the overlap between their own information and the model's.[6] A second analysis found that years of experience, subspecialty, and familiarity with AI did not predict who benefited.[7]
This is where the JAMA perspective and I agree. Knowing when to trust AI gets harder as the model improves, because the cases where you should push back get rarer. Their conclusion points toward taking the human out. Mine is that learning when to push back becomes more valuable, and that it can be taught. The economists' model estimated that many more cases would be best handled by the pair if radiologists correctly combined AI predictions with their own information.[6] That should be the goal. We do not have the evidence to declare it unachievable.
Working as part of a physician-AI team requires a second skill: using AI as a thinking partner. For most doctors the back-and-forth is already familiar. You catch a partner in the hallway, lay out the patient, hear something you had not considered, push back, and one of you changes your mind. We are comfortable ignoring someone else's plan when we think it is wrong. On a good day, we are also comfortable being told our plan is wrong.
How to work with AI
The best description I have found of how this actually goes comes from a field experiment on 758 consultants at BCG, run by a Harvard group in 2023.[8] On tasks within GPT-4's capabilities, AI users were faster and produced better work. On a task outside its capabilities, they were 19 percentage points less likely to get the answer right.[8] The researchers called this the jagged frontier: a model can be excellent at one task and wrong at another that looks much the same.
A follow-up analysis identified three ways consultants worked with the AI.[9] “Centaurs” owned the problem, did the thinking themselves, and handed the model specific pieces. They had the highest accuracy. “Cyborgs” wove the model into the work, asked it to explain its logic, and pointed out contradictions. “Self-automators” handed over the task and accepted what came back. They were less accurate than the centaurs.
There is a warning inside the cyborg result. Some who asked the model to check its own analysis were persuaded into the wrong answer anyway.[9] Arguing with the model is not protective on its own. You need enough expertise to recognize a problem, and something beyond the model's reassurance to check the answer against.
The skilling findings matter too. Centaurs developed more domain expertise, cyborgs developed more AI skill, and self-automators developed neither.[9] That is a warning about learning opportunities that disappear when we stop doing the work.
These are consultants, not physicians, but the result points at a gap in the medical research. In the Goh trials, physicians given AI chose how to use it.[3][4] A later analysis found no significant performance difference between input styles, such as pasting a whole case versus asking short questions.[10] But input style is not mode. Pasting a whole case tells you nothing about whether the physician challenged the answer, checked it independently, or accepted it.
“Physician plus ChatGPT” describes several very different ways of working. Averaging them together may be hiding exactly the difference we need to measure.
What to do next
Keep your clinical expertise sharp. This sounds obvious, and it is the part I expect people to skip. To push back on an answer you have to know enough to see where it is wrong. The deskilling risk is an argument for the reps, not against them.
Learn the back-and-forth, and accept what it costs. Working with a model means admitting you might not know the most about this particular question. We already do this with colleagues we respect. Dismissing the model and deferring to it are both ways of avoiding the harder work of evaluating its answer.
Map the frontier in your own practice. You cannot learn to argue with an output you never asked for. Give the model a real problem from your clinic this week. Push it. Find where it stops being useful, and notice that the edge is not where you assumed, and that it moves.
Do the research differently. Randomizing physicians to access is a solved question with an uninteresting answer. Randomize the mode of use. Put the physician in the room with the patient, and test whether information gathering and deliberate collaboration change outcomes. Then we can say what a physician adds, and what a physician who has learned to work with the model adds on top of that.
The answer is the easy part now. Knowing what to ask, what needs checking, and when to change the plan still takes work.
Nobody is going to hand you those skills. Go try it.