Talk to us
Whether you're buying, selling, partnering, or investing — pick what fits and our team will get back to you within one business day.
A real human, fast
Someone on our team replies within one business day — no bots, no ticket queue.
Routed to the right team
Buying, selling, partnering, or investing — you reach the people who can actually help.
Independent & unbiased
No pushy sales. Just honest guidance grounded in the ecosystem.
Tailored to your context
Tell us what you need and we shape the next steps around it.
Who are you? Pick the option that fits best.
MIT found 95% of enterprise AI pilots deliver zero measurable P&L impact. The gap isn’t the models — it’s everything around them. Here’s what the 5% do differently.
The short version: A 2025 MIT study found that about 95% of enterprise generative-AI pilots delivered no measurable profit-and-loss impact. The cause is not the models — it is the wiring around them: bad data, shallow integration, vague success metrics, and pilots aimed at demos instead of real work. The 5% that succeed treat AI as a business change, not a science project. And they buy far more than they build.
Almost every company is "doing AI." Far fewer are getting anything back. That gap — between the money spent and the value returned — is now large enough to have a name, and a number that keeps executives up at night.
In August 2025, MIT’s NANDA initiative published "The GenAI Divide: State of AI in Business 2025." Drawing on roughly 150 leader interviews, 350 employee surveys, and an analysis of 300 public AI deployments, it reached a blunt conclusion: around 95% of generative-AI pilots produced no measurable P&L impact. Only about 5% were extracting real value — in some cases, millions.
The eye-watering part is the denominator. MIT estimated enterprises had already spent $30–$40 billion on generative AI. Most of that spend, by the report’s measure, returned nothing you could put on a balance sheet. That is the "GenAI Divide": a small group pulling away with compounding advantage while the majority stall in permanent pilot mode.
One caveat, stated plainly: "no measurable P&L impact" is not the same as "useless." Plenty of pilots produced soft wins — morale, learning, faster drafts — that never showed up in the numbers. But if the goal was ROI, the scoreboard is brutal.
This is the most misread part of the story. MIT was explicit that model quality was rarely the bottleneck. The models are good enough. What breaks is everything a demo conveniently skips.
The report’s central idea is a "learning gap." Most pilots use tools that do not remember context, do not adapt to feedback, and do not improve with use. A slick demo handles the happy path; real operations are all edge cases. When the tool cannot learn the messy specifics of how your team actually works, it plateaus — impressive on day one, ignored by day thirty.
Budgets flow to visible, front-office AI — marketing copy, sales assistants, chatbots — because that is where the demos dazzle. MIT found the durable ROI more often sat in the back office: process automation, finance, operations, support triage. Unglamorous, friction-heavy, and exactly where automation compounds. Companies chase the shiny use case and skip the profitable one.
A pilot proves a model can do a task. Production requires integration, permissions, monitoring, change management, and someone who owns the outcome. Most organizations underfund that second, harder half. The result is a graveyard of promising pilots that never touch a real system or a real P&L. Our take on the AI productivity paradox covers the individual-level version of the same problem.
The report landed in the second half of 2025 and became one of the most-cited AI-business findings of the year — partly because it punctured a lot of hype at once. Heading into 2026, the pressure has flipped from "are we using AI?" to "can we prove it paid off?" Boards are asking for ROI, not activity. That makes the GenAI Divide the defining enterprise-AI question of the year: which side of it are you on, and can you show the math? For the bigger structural shift, see whether AI agents will replace your SaaS apps.
Here is the finding most relevant to anyone approving budget. MIT reported that AI tools bought from specialized vendors succeeded about 67% of the time, while internal builds worked only about a third as often — a build-success rate closer to 20–33%. In other words, the confident "we’ll just build it ourselves" instinct was roughly twice as likely to fail.
Why? Successful buyers treated vendors less like a software subscription and more like a service partner — demanding customization, judging tools on measurable outcomes, and insisting they slot into existing workflows. The lesson is not "always buy." It is that for non-differentiating work, a focused vendor with domain fluency usually beats a from-scratch internal project. Our build-vs-buy decision framework breaks down when each path wins.
This part is opinion. The 95% number will be quoted for years as proof that "AI is overhyped." That is the wrong lesson. The right one is that AI is a change-management problem wearing a technology costume. The 1% will stop running science experiments, aim at one unglamorous workflow with a real dollar target, buy proven tools instead of building trophies, and measure like a CFO. The technology is ready. Most organizations are not — and that, not the model, is the moat.
According to MIT’s NANDA initiative report "The GenAI Divide: State of AI in Business 2025," about 95% of enterprise generative-AI pilots delivered no measurable profit-and-loss impact. The main reason was not model quality but the "learning gap" — brittle tools that don’t adapt, weak integration with real workflows, and pilots aimed at flashy front-office demos instead of the back-office processes where the savings actually are.
No. MIT was explicit that the models themselves are largely capable. The failures come from everything around the model: data quality, workflow integration, change management, unclear success metrics, and tools that cannot learn from feedback or remember context. It is an organizational and integration problem, not a model problem.
The 5% that reached real value tended to aim AI at specific, high-friction back-office workflows, define measurable outcomes up front, integrate deeply into existing systems, and partner closely with specialized vendors rather than treating AI as a science experiment. MIT found that tools bought from focused vendors succeeded roughly 67% of the time, versus internal builds that worked only about a third as often.
MIT’s data leaned strongly toward buying: purchased tools from specialized vendors reached deployment about 67% of the time, while internal builds succeeded only roughly one-third as often. Buying is not automatically better, but for most non-differentiating workflows, a proven vendor with domain expertise beats a from-scratch internal build — provided you treat the vendor as a partner and hold them to measurable outcomes.
MIT estimated enterprises had poured roughly $30–$40 billion into generative AI, yet about 95% of organizations reported no measurable return. The gap between spend and results is the core of what the report calls the "GenAI Divide."
The productivity paradox is about individuals and teams not feeling faster despite using AI. The 95%-pilot-failure finding is about organizations: pilots that never cross from demo to production P&L impact. They are related symptoms of the same rollout problem, but one is measured in personal output and the other in business results.
Tags
The 1% Stack
Saaskart's media & intelligence series for software buyers, founders, and operators — opinionated takes on SaaS, AI agents, and the stacks that separate the 1% from everyone else.
Explore thousands of vetted tools, AI agents, and service providers on Saaskart — compare features, pricing, and real buyer reviews in one place.