What Claude is bad at, and which limitations are permanent
An honest inventory of five things Claude will not do well for you, why two of them follow directly from how the technology works, and which ones the next release will not fix.
The list of Claude AI limitations is short, and it repeats. Five of them come up in nearly every session I run, which makes them far easier to write down than the strengths. None is a secret, and knowing them is most of what separates the people who get steady work out of Claude from the people who get lucky now and then.
Two that follow from how it works
The first is current fact. A model knows what it absorbed during training, and it has no live view of the world after that. Ask what a competitor announced last week and, unless you have pasted the announcement in or it has gone and looked, you are inviting it to guess. It will guess fluently, which is the dangerous part.
Arithmetic comes second, at any scale. It produces text by working out what should come next, not by calculating, so a column of figures gets an answer that looks right rather than one that has been added up. Small sums, large sums, percentages in a board paper: check all of them. No release will fix this. Adding up is not what the machine does. Producing likely-looking text is, and a column of figures is text.
Three that follow from what it is not
The third is your organisation. It knows nothing about your pricing, your history with a difficult client, the reason the last restructure went the way it did, or which two directors do not speak. It knows what you have told it in this conversation and nothing else, which is why the same prompt produces something generic for you and something sharp for the colleague who supplied the background. None of that context arrives by osmosis, and supplying it is work somebody has to do.
Then there is judgement that lands on real people. Who to make redundant, whether to escalate a safeguarding concern, which supplier to cut loose, whether a performance case is fair. Claude will produce a considered-sounding answer to all of them, and that fluency is exactly why you should not ask. Somebody has to own the decision, and it cannot be a model. The test I use is simple: if the outcome would be defended in front of the person affected, a human makes it.
Last on the list is being the only reader of anything that matters. Anthropic's own guidance concedes that the techniques which reduce errors do not eliminate them, and the mechanics of a confident wrong answer explain why that concession is structural. If a document is going to a regulator, a client or a court, a person reads it before it leaves.
The full list of jobs to keep away from it altogether runs further than these five. It is best read next to the work it does reliably well, because a boundary only makes sense from both sides.
Which of these the next release will fix
Note which way round this falls. The two that come from how the model works are the two with usable workarounds: give it a live search or paste the announcement in, and put the figures through a spreadsheet. Nothing about the model has changed in either case, you have simply stopped asking it to do what it cannot do. The other three no release touches. No model will ever know what happened in your leadership meeting unless somebody tells it, no update will make it accountable for a decision, and no version number will make it acceptable for a document to go out unread.
Treating all five as temporary gaps is how organisations end up disappointed on a schedule. The productive move is to design around them, and to stop budgeting for the day they go away.
Design for the permanent limits and the rest look after themselves. Design for the temporary ones and you will be having this conversation again in eighteen months, with a larger licence bill.