In July 2024, Gartner predicted that organizations would abandon at least 30% of generative AI projects after a proof of concept by the end of 2025. In January 2026, it reported the real figure: at least 50%. In both cases, it was blamed on poor data quality, inadequate risk controls, escalating costs, and unclear business value. None of these reasons stem from a specific model or model issues, and any developer or technology executive can dig into the codebase or roadmap documents and find all four before the first sprint even starts.
From my perspective, the roadmaps themselves usually look fine. Nobody checks, before the board signs off, whether the company can deliver what’s on them, and that’s exactly what AI implementation readiness is all about. In my experience, it only works when you do it first and when “not yet” is an answer that is allowed to be given to the team.
Enterprise AI roadmap built on assumptions
In the programs I’ve seen, the cost shows up at least six months after approval, which is roughly how long it takes a pilot to reach production and start running into real data. The roadmaps I review tend to list use cases by quarter with a projected saving next to each, and they almost never say whether the data or integrations behind them actually exist. I understand why this is the case, as boards want a plan, and nobody wants to be the person in the room saying “the data isn’t ready.”
Known gaps get the strangest treatment: a missing data pipeline or an ERP with no usable API. These gaps can block every use case that depends on them. On the roadmap, each one turns into one more line item, something like “Q3: integrate pricing data,” sitting in the same column as the features that can’t run without it. And it stays green in the status report right up until the week it blocks something. By then, the company has paid for the build and integration work, which is far more than the assessment would have cost. In October 2024, Forrester predicted that most enterprises fixated on AI ROI would “scale back prematurely,” and a year later it said enterprises would push 25% of their planned AI spend into 2027. I suspect a good chunk of that money is the price of assessments (that no one performed). When the promised returns don’t show up on time, the budget gets cut, and the cut doesn’t care which use cases could (or would) have worked.
We got one of those calls from a mid-sized B2B distributor about a year after its board approved an AI roadmap. The roadmap had come out of six weeks of leadership interviews and listed four use cases: a customer service assistant, AI-drafted quotes, demand forecasting, and supplier risk scoring. As far as I could tell, nobody had looked at the data behind any of them. The service assistant launched first and did OK. Quoting went live a few months later, and that’s where things started going in the wrong direction. The quoting pilot had run on a product catalog export. Unfortunately, someone on the project team implemented that export and refreshed it by hand. In production, contract prices came partly from their ERP system and partly from spreadsheets Sales Ops uploaded every Friday. To be fair, “ERP pricing integration” was on the roadmap. However, it was scheduled for six months after the quoting feature was set to go live, even though the feature needed it to work.
The first real sign of trouble was an email from procurement at the distributor’s largest account. The tool had quoted them list price instead of their contract price, and the rep had sent the quote without reviewing it. I can’t really blame anyone since the drafts had been right for months. Procurement wanted to know, in writing, who had approved that price. Nobody could say because nobody had ever decided who owns an AI-generated price. The distributor issued a credit, switched quoting off, and in the same budget meeting cut demand forecasting too, even though forecasting had nothing to do with it. When we reviewed the project documents, all three problems were visible in the first two weeks: stale pricing data, a missing integration, and no owner for anything the tool might send to customers. Nobody had been asked to look, and a good accounting person would have spotted this in no time.
AI adoption readiness assessments
Some companies do run a readiness assessment first, and it still doesn’t change anything. I’ve written a couple of those assessments myself, earlier in my career, and I thought they were thorough at the time.
Usually the AI maturity assessment starts after the board approves the roadmap, so without anyone saying it out loud, its job becomes confirming the plan. The team that wants the plan approved does most of the scoring, and when other teams participate, the report averages their scores. If the data engineer who cleans up the Friday pricing spreadsheet gives data readiness a value of 2, and the business sponsor gives it a 4, the report says 3. Nobody asks why two people who know the business see the data so differently, even though that disagreement is the most interesting part of the document. RAND’s interviews with 65 experienced data scientists and engineers found that misunderstandings about a project’s purpose are the most common reason AI projects fail, which is exactly the kind of thing that’s hidden by an averaged score.
Fast-forward to today, we collect scores for each function separately and treat any gap of two points or more as a finding to explore and investigate. What comes out is a “go,” “not yet,” or “stop” for each use case, with a named person next to every gap. Otherwise, a recommendation like “improve data quality” with nobody’s name on it just sits in the risk column forever.
We also rerun it (and in some cases, multiple times). Vendors ship new model versions, and the one person who understands the pricing data wins the lottery or moves to a different job. An assessment from March may describe a different company by September, so we repeat ours for each use case before every release and whenever a data source, a model, or an owner changes for any reason.
Two ways to run the same AI readiness assessment:
| The kind that changes nothing | The kind that changes the roadmap | |
| When it runs | After the board approves the roadmap | Before any use case is scheduled |
| Who scores it | The team that wants the roadmap approved | Each function on its own |
| What counts as a finding | The average score | Disagreements between functions |
| What it produces | A maturity score | Go, not yet, or stop for each use case, with an owner per gap |
| How often it runs | Once | Before each release, and when data, models, or owners change |
Assessment coverage
The first thing we look at is the data, and we look at it the way the production environment will see it. In practice, we pull it from the same systems the finished tool will use, as fresh as the tool will get it, messy records and all. A sample that someone cleaned up for the pilot doesn’t tell you much. At the distributor, pulling contract prices the way the quoting tool did would have turned up the Friday spreadsheets in an afternoon. We also look at enterprise AI governance: who can use the data, for what purpose, and under which contracts.
Then there’s the outcome. For every use case, we want today’s number, the target, the date, and the business person who will report on it. That person can’t be the project manager or a steering committee. If we can’t get a number and a name, the use case goes back to discovery. Someone is usually disappointed when that happens, but it’s a lot cheaper than finding out in production.
Workflow is the part people tend to skip. Dropping AI into your existing process is easier to plan, but harder to live with afterward. At the distributor, the quote review step was designed for a rep writing a few quotes a day, and nobody rethought it once a tool could produce dozens of quotes. McKinsey’s State of AI 2026 survey found that nearly three-quarters of AI high performers had fundamentally redesigned workflows because of AI, compared with about a quarter of other respondents. If a use case needs a redesign, we budget for it up front, because anything labeled “phase two” tends to be the first thing cut during budget meetings.
Before anyone writes code, someone has to decide who answers for what the system tells customers, where a person signs off on its output, and ultimately, who can switch it off. Air Canada learned this the hard way in 2024, when it argued that it wasn’t responsible for wrong information its chatbot gave a customer. The British Columbia tribunal called that “a remarkable submission” and held the airline responsible for everything on its website, including the chatbot.
What sounds fine in a workshop versus what holds up in production:
| What sounds fine in a workshop | What we ask to see before production | |
| Data | “We have five years of order history.” | Data from the same systems the tool will use, and who maintains it |
| Integration | “The ERP has an API.” | The integration tested with real volumes, and a plan for when it breaks |
| Outcome | “Faster quote turnaround.” | Today’s number, the target, the date, and who reports on it |
| Workflow | “Reps will review the drafts.” | A redesigned review step with time set aside for it |
| Accountability | “We’ll loop in legal.” | One named person who approves go-live and can switch the system off |
Takeaways
A roadmap is a downstream document, and it can’t be better than the assessment behind it. If nobody took an honest look at the data and the product owners before scheduling use cases, no amount of project management will fix it. The problems will also show up months later, when they cost a lot more to fix. My advice is to do the following 3 things:
- do the assessment before you sequence anything
- let people who don’t own the roadmap do the scoring
- treat “not yet” as a real answer
If your roadmap is already approved, go through it one use case at a time. Also, ensure you ask which assessment finding put each one there and who signed off on it. If the answer is “it was in the AI deployment strategy deck,” then you don’t have a roadmap. Instead, what you really have is a list of hopeful features and/or capabilities, with some attached dates that look realistic.