enterprise GenAI pilots with no measurable P&L impact
MIT Project NANDA studied roughly 300 deployments alongside 52 case studies and 153 leadership surveys.
AI Strategy / June 14, 2026 / 14 min read
Vendor influence is one of the most underexamined reasons AI programmes succeed or fail
By Mohammed Ibrahim
Executive summary
So what
The buyer has to keep authorship of the strategy, the problem, the metrics and the data, then put every vendor underneath an outcome the institution defined before procurement began.
Evidence map
MIT Project NANDA studied roughly 300 deployments alongside 52 case studies and 153 leadership surveys.
RAND's finding came from structured interviews with 65 data scientists and engineers.
A 2025 survey found workflow redesign is the factor most strongly correlated with EBIT impact.
The counterpoint matters
External vendor solutions succeeded about twice as often as internal builds in the MIT evidence. The distinction is between buying execution capability and outsourcing strategic judgement.
Relative success rate
Vendor-partnered approaches
near two-thirds
Internal builds
roughly a third
Ask why enterprise AI keeps failing and you tend to get the same three answers: the technology is immature, the data is messy, the talent is scarce. All three are real enough. They are not the whole story, and probably not the part worth the most attention.
The part that gets the least attention is who decides what gets built, how success is measured, and what gets bought. In a surprising number of programmes those decisions are shaped, and sometimes made outright, by people whose revenue depends on the answer being "more AI, and buy it from us." That is vendor influence, and it barely surfaces in the post-mortems.
The case here runs along three lines. The failure rates are real, large, and consistent across independent studies. The causes are mostly organisational rather than technical, and the most common single one is that the wrong problem got chosen at the start. And vendor influence is not a bad thing in itself; it turns corrosive at one identifiable point, where a vendor stops supplying capability and starts authoring your strategy. Cross that point and you have handed off the decision that most determines whether AI works at all, what problem to solve and who owns the result, to a party with a commercial stake in getting it wrong.
None of this adds up to "avoid vendors." The best evidence points the other way. Working with specialised vendors roughly doubles the odds of success against building everything in-house. The point is narrower and harder than that: the buyer has to keep authorship of the strategy, the problem, the metrics, and the data, and treat vendor influence as the concentrated risk it plainly is. Most do not, and that is where programmes come apart, more reliably than any limitation of the models. The whole of it compresses to six words.
Buy the tool,
own the thinking
01 Evidence
It is worth establishing how much failure there is to explain before getting to why.
MIT's Project NANDA studied roughly 300 deployments alongside 52 case studies and 153 leadership surveys. Around 95% of enterprise generative-AI pilots delivered no measurable impact on profit and loss, and only about 5% of integrated systems produced significant value.¹ The authors located the cause not in model quality or regulation, but in approach.
RAND reached a similar place from a different method. Working from structured interviews with 65 data scientists and engineers, it put the failure rate above 80%, about twice the rate of conventional IT projects.² Its headline finding is the one that should be pinned to a wall somewhere: the most common cause of failure is not the technology, it is that leaders misunderstand or miscommunicate the problem the AI is meant to solve.
A 2025 survey of nearly 2,000 respondents picked up the same paradox from the value side. Most organisations now use AI somewhere, but only a small minority can attribute real EBIT impact to it, and the factor that correlates most strongly with that impact is fundamental workflow redesign, which only about a fifth of adopters actually attempt.³ The rest are laying AI over processes they never rethought.
Read across the three and the shape is consistent. Adoption is close to universal; value capture is rare and clustered in a few hands. The bottleneck was never access to AI. It is the discipline to point it at a problem worth solving and wire it into how the work actually flows.
02 Market story
The comfortable explanation is that AI is early. The models will improve, the data will mature, so be patient, keep buying, and the value will arrive.
It is comfortable because it lets everyone off the hook. The buyer's pilot failed because "the technology wasn't ready." The vendor sold a fine product to a customer who "wasn't mature enough to use it." And the budget stays open for next year's pilot, which will begin life sitting on the same unresolved data and ownership problems as the last one.
The evidence does not really support that story. RAND's root causes are about people and choices far more than silicon: stakeholders misread the problem, the data is not usable, the newest technology gets chased instead of a real problem getting solved, the infrastructure to run and maintain models is missing. Four of the five are failures of judgement and governance.
That is the pivot the whole argument turns on. If failure is mostly a matter of problem selection, measurement, and integration, the question that matters is not "which model" but "who is in the room when the problem gets chosen, and what do they need to be true?" In a great many programmes, the loudest and best-funded voice in that room belongs to a vendor.
03 Vendor influence
There is nothing conspiratorial about this, and it does not require anyone to act in bad faith. It is what you get when misaligned incentives run through ordinary commercial behaviour. Six mechanisms do most of the work.
The incentive gap. A buyer succeeds when AI moves the P&L. A model provider succeeds when you consume more tokens. A platform succeeds when more of your workloads run on its stack. A licensed-software vendor succeeds when seats are provisioned, used or not. An integrator on time-and-materials succeeds when the engagement is large and runs long. None of those is the buyer's success metric. They track it in good conditions and pull away from it in bad ones, and the bad conditions are exactly the moments when a vendor's advice carries the most weight. When a project is in trouble, the buyer's interest is to stop and reassess; almost every vendor's interest is to widen the scope and press on. The advice you get at the point of maximum doubt is the advice most distorted by this gap.
Solution-first selling. Vendors increasingly position AI as a default feature rather than a deliberate one, bolting models onto products where they add cost and risk surface rather than value. The buyer-side version is the executive who opens with "we need an AI strategy" or "we need agents" instead of "we have this expensive, recurring, well-understood problem." Start from the technology and the problem gets quietly reverse-engineered to fit whatever is for sale.
Proof-of-concept theatre. The pilot is where the two sides' incentives line up just long enough to be dangerous. It is cheap, fast, and almost always succeeds, because it was built to. It runs on a hand-picked slice of clean data and skips the integration, the governance, and the change work that make up most of the real job. It produces a slide that says it works, and then the project settles into pilot purgatory. The pilot was never a test of value; it was a sales instrument the buyer mistook for evidence.
AI washing. As AI became the default pitch, the gap between claimed and actual capability widened into what is now openly called AI washing, and regulators have started treating the worst of it as deceptive marketing. The supply-side problem is worth saying plainly: a wave of firms rebranded as AI experts without the track record underneath. The awkward part for a buyer is that the people most confident about what AI can do are often the people selling it, and confidence is not competence.
Lock-in. Even where AI works, vendor influence shapes how much of the value the buyer gets to keep. Lock-in runs through data gravity, since moving large datasets is expensive and egress alone is often 10 to 15% of a cloud bill. It runs through proprietary runtimes and orchestration layers, and through the plain inertia of an integrated stack. None of that is a reason to avoid vendors. It is a reason to architect for portability and to price the dependence into the decision up front, which a vendor will seldom volunteer to help you do.
The integrator conflict. The consulting and integration layer deserves a harder look, and as an independent firm we hold ourselves to the same test. Two structures bend the advice. Reseller economics: many integrators earn their margin and partner status by moving a particular vendor's product, which quietly tilts the supposedly neutral recommendation. And body-shop economics: when the revenue model is billable hours, the pull is towards bigger teams and longer timelines rather than the smallest intervention that would actually solve the problem.
Influence mechanics
None require bad faith. They become dangerous when commercial incentives start shaping problem selection, measurement and continuation
The buyer needs P&L movement. The supplier may be paid for tokens, seats, workloads, margin or hours.
Technology arrives first, and the problem is reverse-engineered to fit what is already for sale.
A clean-data pilot becomes a sales instrument the buyer mistakes for production evidence.
Confidence and branding can outrun track record, especially when AI becomes the default pitch.
Data gravity, egress, runtimes and orchestration layers shape how much value the buyer keeps.
Reseller economics and body-shop economics can bend neutral advice towards bigger or longer engagements.
04 Regional lens
For organisations in the Middle East, vendor influence carries an extra charge, because the dependence is playing out at the enterprise level and the national level at once, and the two feed each other.
The ambition is genuine and large. Generative AI is projected to contribute on the order of USD 23 billion a year across the GCC by 2030, roughly 2% of regional GDP, and Gulf states had committed well over USD 30 billion to AI by early 2025.⁶ Saudi Arabia named 2026 its year of AI; the UAE is anchoring some of the largest compute build-outs anywhere. The instinct underneath the spending is a sound one: never again build an economy on infrastructure you do not own.
The operating picture is more sober, and it rhymes with the global data. Most regional organisations are using AI, but few are running it at scale or capturing value, the same pilot-bound pattern as everywhere else. Regional research documents a widening gap between corporate ambition and operational readiness, with shortages of local specialists and thin strategic-planning capacity.⁷
Three things follow for Gulf buyers in particular. The first is that the sovereignty paradox is real: the region is investing, rightly, in owning its AI infrastructure, yet the quickest route to capacity runs straight through foreign hyperscalers, chipmakers, and model providers. Sovereign ambition and vendor dependence are, for now, two sides of one contract, and the task is to capture the capability without surrendering the control. The second is that the readiness gap makes the region more exposed to vendor influence, not less. Where local strategy and data capability are thin, the vendor's voice fills the vacuum, and the vendor is glad to supply not just the tool but the roadmap and the definition of success along with it. That is the strongest argument for building independent, incentive-aligned advisory capacity that sits on the buyer's side of the table.
The third is that the exposure is not spread evenly. The largest institutions, the tier-one banks and the sovereign-backed champions, are building internal capability and can increasingly meet vendors as equals. The harder position belongs to the tier-two banks, the insurers, the family conglomerates, and the mid-market firms that have the ambition and the budget but not yet the in-house muscle to keep a sophisticated vendor honest. For that group the foundational work, strategy and data and governance and change readiness, is not a precondition for the AI programme. It is the AI programme. Everything else is procurement.
Gulf pressure
projected annual GCC GenAI contribution by 2030
approximate share of regional GDP
AI commitments by early 2025
05 Counterpoint
An honest case has to deal with its best counter-evidence, so here is the strongest piece of it. The same MIT study behind the 95% headline also found that AI solutions built or supplied by external vendors succeeded about twice as often as those built in-house: a success rate near two-thirds for vendor-partnered approaches, against roughly a third for internal builds.⁴
So if vendors are as corrosive as all that, why do their solutions win more often?
Because "vendor influence over strategy" and "vendor capability in execution" are two different things, and the figure measures the second. Specialised vendors bring proven tooling, tuned models, and implementation patterns an internal team starting cold usually cannot match. Buying a fit-for-purpose tool from a firm that has solved your problem fifty times is the smart move. It is the opposite of the failure mode.
The failure mode is letting the vendor's commercial logic decide which problem you solve, how you measure whether it worked, what you buy to solve it, and whether you keep going when the evidence says stop. You can buy the capability and still refuse to outsource the judgement.
So the answer is not "build it yourself," which is the worse-performing path. It is shorter than that.
Buy the tool. Own the thinking.
Keep authorship of the problem, the metrics, the data strategy, and the build-versus-buy call itself, and put every vendor, however good, underneath a problem and an outcome you defined before any of them walked in. The organisations in the successful 5% tend to treat vendors the way a sophisticated company treats an outsourcer: they demand customisation, they measure business outcomes rather than technical specifications, and they keep their hand on the relationship throughout.
06 Success profile
If the failures cluster around problem selection, measurement, integration, and ownership, the successes cluster around the reverse, and the same profile keeps turning up across the evidence.
They tend to start from a specific, expensive, recurring problem, and usually an unglamorous one. MIT found returns are highest in the back office, in finance, operations, and compliance, rather than in the sales and marketing functions that absorb more than half of AI budgets and return the least.⁵ They redesign the workflow instead of decorating it, which matters because the 2025 survey's single strongest correlate of EBIT impact is exactly that redesign, and only about a fifth of adopters get to it. Bolt a copilot onto a broken process and you get a faster broken process.
They also do the boring foundational work first. Usable, governed data; clear ownership; infrastructure that exists in reality and not just on a roadmap. The "next pilot, same data problem" loop is the tell of an organisation that keeps skipping this step. And they keep the vendor on a short leash: outcome-based measurement, portability designed in from the start, and a problem they defined themselves.
It is the same conclusion the failure researchers reach from the opposite direction. The gap between adoption and value is a leadership and operating-model problem, not a technology one. For most enterprise uses the technology is already good enough. The discipline is the scarce part.
07 Buyer playbook
If vendor influence is a concentrated risk, it is worth governing like one. Seven habits keep the capability while keeping the judgement.
Define the problem before you meet a vendor. Write down the specific, costed, recurring problem and the metric that would prove success, in business terms (cost, revenue, cycle time, error rate) and never in technical ones (accuracy, model size). If a vendor helped you write it, start again.
Separate the advisor from the supplier. The party that helps you decide what to do, and whether it worked, should not be the party that profits from the what. Where they have to be the same, make the conflict explicit and price it in.
Treat the pilot as a sales instrument until it proves otherwise. Insist it runs on representative rather than curated data, and that it carries the integration, governance, and change work that decide whether it survives production. A demo is not evidence.
Measure outcomes, not activity. Tie payment and continuation to business results. Outcome-based or value-share commercial models push risk back towards the party best placed to manage it, and they quickly reveal which vendors believe their own pitch.
Architect for exit from day one. Assume you will want to leave. Quantify egress, data portability, and re-platforming costs before you sign, and prefer open standards and model-agnostic layers wherever the economics allow. Dependence you chose with open eyes is a strategy; dependence you discover later is a trap.
Fix the foundations on your own account. Data quality, governance, ownership, and workflow redesign are buyer-side jobs no vendor can do for you, or is incentivised to. They are also the highest-correlated drivers of value in every serious study, so spend here first.
Be willing to stop. The most valuable and least rewarded act in any AI programme is killing the one that is not working. Build the off-ramp before you need it, and make sure the person who controls it does not bill by the hour.
Buyer discipline
01
Write down the specific, costed, recurring problem and the metric that would prove success.
08 Close
The usual story blames the technology, partly because the technology cannot answer back and partly because everyone in the value chain benefits from the alibi. The evidence tells a less flattering version. AI programmes fail mostly because the wrong problem was chosen, the workflow was never redesigned, the foundations were never built, and the thing kept running well past the point where it should have been stopped, and each of those is the decision most exposed to a vendor whose idea of success is not the buyer's.
This is not an argument against vendors. Specialised vendors clearly improve the odds, and a buyer who insists on building everything alone is choosing the worse path. It is an argument against abdication, against letting the supply side write the strategy, set the metrics, run the only test, and decide whether to carry on.
The organisations capturing real value have not found better models than everyone else. They have done something less glamorous and a good deal harder: they kept authorship of their own transformation. They worked out what problem mattered, how they would know it had worked, and what they would buy to get there, and then they held even their most capable vendors to that line.
So the question worth asking before the next AI investment is not whether the technology is ready. It is whose transformation this actually is, and who in the room gets paid more if the answer is yes.
The question
It is whose transformation this actually is, and who in the room gets paid more if the answer is yes
What problem matters enough to fund?
How will success be measured in business terms?
Who owns the data and the rules around it?
Who can stop the programme when evidence says stop?
Mohammed Ibrahim is a co-founder of Mal7, focused on enterprise AI strategy, hyperscaler ecosystems and commercial growth for regulated institutions moving AI into production.
Keep authorship of the transformation