There is a particular kind of meeting that happens about four months after an automation project starts. Everyone agrees the tool works. Everyone agrees it is not being used. The conversation goes in circles for forty minutes and then someone suggests more training.
The training does not help, because the tool was never the problem.
What MIT actually found
In 2025, Project NANDA at the MIT Media Lab published The GenAI Divide: State of AI in Business, built from a review of over 300 publicly disclosed AI initiatives, structured interviews with 52 organizations, and survey responses from 153 senior leaders. Its headline is worth quoting exactly, because almost nobody does.
Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.
You will have seen that reported as "95% of AI pilots fail." That is a different claim about a different denominator, and the difference matters if you are trying to work out whether it applies to you.
The more useful sentence is the next one:
This divide does not seem to be driven by model quality or regulation, but seems to be determined by approach.
Not the models. Not the rules. The approach.
The failure is a fit problem
The report tracks what happens to enterprise-grade AI tools as they move through an organization. Sixty percent of organizations evaluated them. Twenty percent reached a pilot. Five percent reached production.
Most of the losses, in the report's words, come down to "brittle workflows, lack of contextual learning, and misalignment with day-to-day operations."
Read that list again. None of it is about capability. It is a description of software that does something impressive in a demo and then meets an actual Tuesday.
One CIO quoted in the report puts it more plainly than any analyst would: "We've seen dozens of demos this year. Maybe one or two are genuinely useful. The rest are wrappers or science projects."
Where the design was supposed to happen
Here is the part the research describes but does not quite name.
Before you can automate a process, that process has to exist as something more than a habit. Somebody has to be able to say what happens, in what order, who is responsible at each point, and what the system should do when the normal path does not apply. Most operations cannot produce that description. Not because they are badly run — because the process lives in the heads of the three people who do it, and it has never needed to be written down.
Automation does not tolerate that. A person encountering an unusual case improvises. A workflow encountering an unusual case does whatever it was told to do, which is usually nothing useful.
So the automation gets built against the version of the process everyone agreed on in the kickoff meeting, which is the tidy version, which is not the one that runs. Writing down the real one is unglamorous and takes an afternoon; here is how to do it. It works for the cases that match. It fails or does something strange on the rest. And after a few weeks of that, people quietly go back to doing it by hand, and nobody logs the moment the project died.
What this predicts
If the failure is about fit rather than capability, a few things follow.
Buying a better tool does not help, because the next tool will also meet the undefined process. Neither does more training, which teaches people to operate something that does not match what they actually do. And a pilot that goes well means less than it looks like: pilots run on the tidy path, with attention on them, usually with the people who care most. The gap between 20% reaching pilot and 5% reaching production is that difference.
It is also worth noticing which automations do survive. They are almost never the ones that demoed well — see nobody demos the boring one.
The thing that does help is unglamorous and comes first. Write down what actually happens — including the exceptions, especially the exceptions. Decide who owns each step. Decide what the system does when the case is unusual, which is more often than anyone expects. Only then does the tooling question become answerable, and by that point it is usually the least interesting decision left.
The uncomfortable version
The reason this pattern survives is that the diagnosis is unflattering to everyone in the room.
A vendor cannot say it, because the fix is not something they sell. An internal sponsor cannot say it, because it means the last two quarters went into a project that skipped its first step. And the team cannot say it, because "our process is not written down anywhere" sounds like an admission rather than a completely normal state of affairs for a business that has been growing.
So the meeting concludes with more training, and everyone goes back to work.
The four-month meeting is not really about the tool. It is the first moment anyone is forced to look at the process, and by then the budget is spent.
Which is an argument for looking earlier, when it costs a week instead of a quarter — and for treating the design of the system as the work rather than as the paperwork before the work.
An AI Systems Review looks at your revenue and operations.
Where work depends on someone remembering, what it costs, and what would change first.
Book an AI Systems Review →