In December I asked Claude to build a powerpoint presentation to go along with a talk.

It was terrible.

That verdict carried weight with me, because I care about presentations. I give a lot of talks, and I am extremely particular about how my slides look and how a story unfolds visually. So, I closed the laptop with what felt like a settled conclusion: AI is not going to make my PowerPoints. I said that to people. I laughed about it. AI was never going to take that work away because my slides are special.

Three months later I was short on time and tried again.

The models had improved. It asked for an example of my prior work, which I gave it. It copied my structure, my style, the way I sequence a story. And it gave me a deck I actually liked. Not passable slides I tolerated because they saved an hour. But 80% of a presentation I could stand up and give. It was wild. For me, it was mind-blowing.

If I had not happened to try again, I would still believe AI could not do this. No one is on Twitter talking about how AI is really great at making a surgical education powerpoint. The companies aren't marketing the new models to me. I would never have known this without just trying it.

My judgment was not wrong. It was expired.

The skill nobody lists first

When physicians ask what they need to learn before using AI, the usual answers are prompting techniques or some working knowledge of how the models are built. That's what I've been telling people, too. But I think the first skill is more basic: learning how to try it.

Take a task you actually do, hand it to an AI system, and find out whether the result is good enough to use in your hands, right now.

Maybe that sounds too simple to call a skill. But it isn't. This is a very specific skill, and not one you learn only once. You have to learn the process of repeating it. That process can be boring, tedious, monotonous. Learning to push through that is what will get you to a breakthrough. So it's absolutely a skill.

With ordinary software, we build durable mental models. If the EHR cannot do something today, it cannot do it tomorrow, and we stop asking. Every experienced clinician carries hundreds of these quiet conclusions about what tools can and cannot do, and for most technology they hold true for years.

With AI they can expire in a quarter. My PowerPoint conclusion lasted three months. Which raises an uncomfortable question: how many things have we each decided AI cannot do, based on an experiment we ran six months ago?

Why your intuition will keep failing you

It gets worse, because AI capability does not follow the curve we expect, where easy tasks succeed and performance degrades smoothly as tasks get harder. To use an oft-cited expression in tech: “The frontier is jagged.” A model can produce a sophisticated statistical critique and then trip on something trivial. (This is the part where the model that just solved an 80-year-old math problem can't tell you how many Rs are in strawberry.) The truth is that human intuition about task difficulty is a poor predictor of model difficulty.

There is decent evidence for this. Dell'Acqua and colleagues ran a field experiment with 758 consultants doing realistic knowledge work. On 18 tasks inside the model's capability frontier, consultants using AI did significantly better and faster. On a task deliberately chosen to sit just outside that frontier, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it. Same people, same tool, opposite effect... and the task outside the frontier did not look harder.

What about “benchmarks”? AI labs publish results on standardized tests, including medical ones, and those numbers matter for comparing models. But a benchmark score cannot tell me whether a model can critique the methodology of a vascular surgery paper without inventing flaws, or turn my messy notes into the structure I actually need, or analyze a quality dataset without a silent statistical error. The benchmark I care about most does not exist. Yours probably doesn't either. I need a Marissa PowerPoint benchmark and that's extremely specific to me.

Physicians already have the most important ingredient

I think physicians misjudge their own position.

Many of us feel behind because we are not AI experts. But the scarce ingredient in all of this is not knowledge of transformers. It is the ability to look at an output and know whether it is any good. That is domain expertise, and we have spent our careers building it. “Deep domain expertise.” Something the tech bros keep pointing out is so necessary. We have it. It's kind of our thing.

A vascular surgeon recognizes a subtly wrong statement about when to intervene on a carotid stenosis. A statistician notices the analysis that should have been stratified. An experienced clinician can tell when a beautifully written assessment does not actually hang together clinically. A novice cannot do any of that, no matter how well they prompt.

Experimentation is the AI skill. Domain expertise is what makes the experiment meaningful.

Which leads to a practical rule: start by testing AI on work you already know how to do. Do not begin by delegating something outside your expertise. That is precisely where you are least equipped to notice failure, and where the consultants in that study got burned.

Build your own tiny benchmark

Here is what trying it looks like when you do it deliberately.

Pick a few representative tasks from your real work. Choose tasks where you know what good looks like, where you would recognize the errors that matter, and where the task recurs often enough that help would change your week.

For each one, decide what success means before you run anything. Suppose the task is critiquing a clinical paper. Success might mean the model finds the real methodological weaknesses, separates major limitations from minor ones, invents nothing, interprets the statistics correctly, and gives criticism specific to the clinical domain rather than generic caution.

Then run it. Run it on more than one model if you can. And decide if the output saves you time or just gives you more work to review.

Stop asking whether AI is good

We spend a lot of energy on questions like whether AI is good at medicine, or research, or writing. Those questions are too broad to have answers. It depends on the model, the task, the context you give it, the standard you hold it to... and on when you last checked.

So pick one task you know cold. Define success before you start. Give it to the best AI tool you have access to, and watch where it succeeds and where it breaks. Map the jagged frontier. Decide where AI will help you and where it's still falling short.

This piece was partly inspired by The AI Daily Brief's episode “The AI Engineering Skills Map for Knowledge Workers” (August 18, 2026), which maps these skills for knowledge workers broadly. I think the physician's version starts here.

Reference: Dell'Acqua F, McFowland E, Mollick E, et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School Working Paper 24-013, 2023; published in Organization Science, 2025.