Why AI projects fail.

A demonstration is judged on whether it impresses a meeting. A working system is judged on whether it survives real customers, real paperwork and real staff on a Monday morning. Those are different tests, and most projects are only ever built to pass the first one.

The short answer

AI projects fail for six recurring reasons: nobody agreed what working would mean, it was built on tidy information that does not exist in your business, only the easy cases were considered, nobody was left holding it afterwards, the attention went to the technology rather than the process around it, and it never reached the people actually doing the job.

None of these are things you could reasonably have known to ask for. They are engineering and ownership problems, which is why buying newer technology almost never rescues a stalled project.

01 The causes

Six reasons it never reached the business.

01 Measurement

Nobody agreed what working would mean

This is the most common cause and the hardest to spot from the outside. Something gets built, it looks good in a meeting, and then somebody asks for a change. Make it shorter. Handle the French contracts. Stop it guessing at account numbers.

The change gets made, and now nobody can answer the only question that matters: is it better than it was this morning? It is different. Whether it is better becomes a matter of opinion, and opinion cannot get software approved for release.

The fix is dull and decisive. Before anything is built, you take a few hundred real cases from your own business, write down the right answer for each, and score every version against them. From then on, improvement is a number rather than an argument.

02 Your records

It was built on tidy information you do not have

Demonstrations are built from a curated sample. Somebody pulls a few hundred neat records and the software handles them beautifully.

Your actual records are not that. They are scanned at an angle, with the customer name typed into the notes field, three date formats in one column, a decade of categories that changed meaning twice, and entries somebody made in 2019 while on the phone. Accuracy collapses, and it does so quietly, because software of this kind does not usually fail loudly. It produces a confident answer that happens to be wrong.

Fixing this is mostly unglamorous tidying rather than anything to do with AI, and it is the single most underestimated part of these projects.

03 Exceptions

Only the easy cases were considered

A demonstration shows the case where everything is in order. Real work is mostly exceptions. The invoice with no purchase order. The claim where the policy lapsed halfway through treatment. The contract in a language nobody planned for.

Today those are handled by somebody who knows the business, uses judgment, and quietly absorbs the mess without anyone writing it down. Because it was never written down, it was never scoped, so the system cannot do it. Your staff find the gaps within a day and stop trusting it immediately.

The remedy is straightforward and usually skipped: sit with the people doing the job and catalogue what they handle that the process documentation does not mention.

04 Ownership

Nobody was left holding it

The project was run by a transformation team, an innovation group, or a consultancy. The demonstration lands, everybody agrees it is promising, and then the question arrives: who owns this now?

Your IT function did not build it and will not adopt something they cannot support. The people who built it have moved to the next client. The department that wanted it has no technical staff. So it sits, and after a quarter or two it stops appearing on the agenda.

This looks like a technology failure and is actually an organisational one. It is also the reason we run what we build rather than handing it over and leaving.

05 Attention

The wrong thing got the attention

Which AI technology to use is the most discussed and least decisive part of these projects. It gets the attention because everyone has an opinion available to them and the names are in the newspaper.

What actually decides whether it works is duller: how the information gets found, what happens when something takes too long, whether the system can admit it does not know, who is allowed to see what, and whether a wrong answer can be caught before somebody acts on it.

A project that spends a month choosing technology and a week on those questions will produce a worse result than one that does the reverse, using exactly the same technology.

06 Adoption

It never reached the people doing the job

The system lives in its own window. Somebody opens a separate tool, pastes something in, reads the result, and copies it back into the system where the work actually happens.

That is acceptable in a demonstration and fatal in a business, because it adds a step to a job people are already too busy to do. Usage stays flat and the project is written off as a technology failure when it was a placement failure.

Systems that get used appear inside the software the person already has open, and the result lands where the next action is taken. This is unglamorous plumbing, and it is usually the difference between a system people use and one they do not.

02 Triage

Can yours be saved?

Usually worth saving
  • The process chosen genuinely costs the business real money
  • It worked on the tidy examples, so the idea is sound
  • What is missing is the finishing work rather than the concept
  • Somebody senior still wants it to exist
Building it again will fail again
  • The rules are fixed and well understood, so ordinary software is the better tool
  • Every answer must be exactly right, with no room for a person to review
  • A confident wrong answer costs more than doing the whole thing by hand
  • The process itself is the problem, and no software will rescue it
Two weeks Roughly how long it takes to tell the two apart, looking at the real system and your real records
Two to four months To take a salvageable process the rest of the way into the business
Not the technology Buying something newer changes the answers without telling you whether they improved
03 Questions

The things people ask next.

Was this our fault?

Almost certainly not. None of the six are things a buyer without a technical background could reasonably have known to ask for, and firms that sell demonstrations have little incentive to raise them. If you want one thing to ask any future supplier, ask how they will prove the system got better after a change.

Do we need to hire technical people to avoid this next time?

No. Every one of these is the responsibility of whoever builds and runs the system. They are not things a business needs to learn, which is why an arrangement where somebody else builds it and keeps running it avoids them without any hiring on your side.

Would newer technology fix it?

Rarely. If nobody can measure the system, newer technology changes the answers without anyone being able to establish whether they improved. Put the measurement in place first, and the technology question answers itself with evidence rather than opinion.

What would it cost to put right?

We publish our prices. A two week assessment starts at $15,000 and tells you whether it is worth doing at all. Putting one process properly into the business starts at $75,000. The pricing page has the rest.