Why your team's Claude results are patchy

Two people swear by it, the rest have an account they opened once, and nobody raises it in a meeting because raising it sounds like a confession about yourself.

Most AI adoption problems in teams look identical from the outside. The licences were signed off months ago. Two or three people are keen and will tell you, unprompted, how good it is. Everyone else opened an account in March, asked it something, got a competent paragraph of nothing, and quietly went back to writing the thing themselves. The results are patchy and nobody says so out loud, because saying so in a meeting sounds like an admission about your own week rather than about the tool.

The evidence points at the operator, not the tool

Anthropic's Economic Index is the most useful thing published on this, because it measures what people actually do, not what they say in a survey. Its central finding for anyone running a team: users registered six months or more are 5% more likely to successfully automate a task than newer users.

Taken alone the number is modest. What makes it matter is that the gap persists after controlling for model type, language, use case and location. Same tool, same job, same country, different outcome. The only variable left standing is the person doing the typing.

Which is why switching platforms so rarely fixes anything. A team getting uneven results from Claude will get uneven results from whatever it moves to next, and will have spent two quarters and a procurement process finding that out.

What the research actually found Six months or more iterating on the task 28.2% The first few months handing over the whole job 38.1% 5% more likely to complete an automated task successfully, and the gap holds after controlling for model, language, use case and location.
The tell is in the second bar. Newer users try to hand over the whole job. Experienced users work the task with it. Same software, different result. Figures from Anthropic's Economic Index, June 2026.

Beginners hand over the whole job. Experienced people iterate.

The same report splits interactions by type, and that split is the heart of it. High-tenure users engage in what the researchers call task iteration 28.2% of the time. They treat Claude as a collaborator: send something, read what comes back, correct it, narrow it, send again.

Low-tenure users issue directive tasks 38.1% of the time. A directive task hands the whole job over in one message and expects it back finished. Write the strategy. Summarise the pack. Draft the redundancy policy. One instruction, one expectation, no second turn.

I can usually tell which group somebody is in from their first prompt of the morning. Directive users send one sentence, read the reply as a verdict on the software, and never think to answer it back, because nobody has told them they are allowed to.

The directive approach is the intuitive one, and it is what most of the marketing has trained people to expect. It also produces flat, faintly generic output, which leads sensible people to conclude the tool is overrated. They have drawn a fair conclusion from an unfair test.

What that difference looks like on the page

A before and after prompt on my homepage shows this in ten seconds. The same request, written twice. The first asks for something about our new product for LinkedIn, and not too long. The second names the voice and the audience, sets the shape and the length, and pastes in two past posts to match.

Nothing in the second version needs technical knowledge. It needs the person to accept that the specification is their job rather than Claude's. That is a habit, not a feature, and habits are what deliberate practice at prompt writing changes.

Why nobody in the room mentions it

Patchy results survive because they are socially awkward. The two enthusiasts are not hoarding a secret. Ask them what they do differently and most cannot tell you, which is a problem in itself, since what cannot be described cannot be copied. Meanwhile the sceptics have a private theory that the whole thing is overblown, and one bad experiment is enough to keep it.

If that sounds like your team, work out whether the constraint is skill before you buy anything else. Four questions will tell you whether training would move it. If you are the one who signed for the licences, the founder-led version of this problem has a shape of its own.

Those two enthusiasts will move on eventually, or get promoted, and the capability leaves the building with them. That is the version of this problem that reaches a P&L, and it reaches it late.

Sources

  1. Anthropic — Economic Index, June 2026 report
  2. Built In — Anthropic's Economic Index and the 2026 AI jobs report