The Skill to Trusting AI

AI just handed you something like a draft, an analysis, or a plan which cost almost nothing to make. Now you have to consider what it costs to trust. Every piece of machine output still leaves someone with the job of knowing whether it is right. When that responsibility gets passed downstream, the person receiving the work inherits something they didn’t create and may not have the context to evaluate.

And checking work you didn’t do is harder than it sounds. This is counterintuitive, because generation feels like the difficult part and checking feels like the easy part. But when you build an analysis yourself, you walk the causal chain as you go. You know why each step is there, what assumptions you made, and where something might have gone wrong. When a machine builds it, you have to reason backwards from the answer to reconstruct enough of that chain to decide whether you trust it. Those are different cognitive tasks, and we tend to underestimate the second one.

METR actually measured a piece of this. They ran a randomized trial with 16 experienced developers working on their own codebases—repositories they had worked in for years. Beforehand, the developers expected AI to make them about 24 percent faster. After doing the work, they still believed it had made them 20 percent faster. In reality, they were 19 percent slower.  

Part of the time went into the work around generation: prompting, waiting, reviewing and editing. The developers accepted fewer than 44 percent of the AI-generated code, and about 9 percent of their total task time went to reviewing and cleaning up AI output.  

The interesting thing is that the generation was visible and the verification was not. Code appeared almost instantly, which made the work felt faster. The time spent figuring out whether that code was actually right was distributed across the task, and people didn’t seem to experience it as the same kind of work.

That matters far beyond software. As AI moves further into knowledge work, we are going to do more work by evaluating things we didn’t produce ourselves. And that changes the value of expertise: you need enough understanding to know not just whether an answer looks plausible, but whether it is actually right.

A second study helps explain the not-noticing. Hu, Cao and Li surveyed 544 AI users and found that the more people trusted the system, the less they tended to verify its outputs. Trust goes up, checking goes down. They also found that the things that make AI more appealing—its versatility, creativity and ability to work across different kinds of media—can increase that trust.  

Which means the unpaid debt can pile up fastest in exactly the places where people feel most comfortable with the system.

This is systemic. The front line approves machine drafts. Managers approve the approvers, usually with less context than the approvers had. And at the top someone signs—a forecast, a filing, a key decision—resting on chains of machine work that nobody in the building has walked end to end.

So what is the checking, actually? Four questions. Each one costs more than the last and each one requires different kinds of knowledge and responsibilities.

  • Is there an error?

  • Is it right?

  • Will it work?

  • Do I trust it enough to act?

Finding an error is forensic. Machines can help with that one, honestly. Knowing it’s right takes expertise—an analysis can be mistake-free and still misread the customer completely. Knowing it’ll work means understanding how this will behave in the real world: people, timing, incentives, all the organizational politics. And trust is about gets squandered if you’re wrong.

Most companies are staffed, implicitly, for question one. The exposure is on three and four. A board’s entire job is largely question four.

Here's a radical idea in the AI age: I think checking is where people can grow now.

An analyst who traces a claim back to its source is building something in herself. A manager who asks, “Okay, but would this survive contact with an actual customer?” is teaching everyone in the room how to think. An executive who refuses to sign something nobody can defend is making the standard for judgment visible to everyone around them.

Something happens to you when you actually ask those four questions of a piece of machine work. Every claim you trace, every assumption you catch, every time you say, “I’m not putting my name on this yet”—judgment is forming.

We used to have a saying in an analytical strategy group I was in: the person who builds the model learns the most. That used to happen inside the making. You made the analysis, wrote the memo, built the model, and your understanding grew along the way. As AI takes over more of the making, we have to be much more deliberate about where that learning happens.

This is why I keep coming back to authorship. The author is the person who can stand up and take a question about the work, whoever—or whatever—produced the first draft. They understand enough to defend it, challenge it, change it, or refuse it.

And this is what I mean when I say stay the author of your own mind. AI is already getting into your process, changing what things mean to you and becoming part of how you think. The question is whether you are directing that process. The checking is one of the places you can.

In our research with more than a thousand people, we kept seeing the same thing: people who used AI heavily while actively directing their relationship with it tended to come out stronger in their work and clearer about their own contribution. People who simply handed things over had a different experience. The work got easier but their relationship to the work got more tenuous.

That makes the practice surprisingly concrete: know what role you are giving to AI, know what you still need to understand yourself, and keep putting your name only on work you can actually stand behind.

We answer these live in the Stay Human Briefing. Your team, your questions.

Click here to see more on the briefing as well as our keynotes and workshops.