Why AI Isn't Delivering Results in Manufacturing (And Why More AI Won't Fix It)
68% of manufacturers have deployed shop-floor AI. Only 16% hit their targets. The gap isn't AI capability, it's what the AI was pointed at.
August 17, 2026 · Updated August 17, 2026
Sixty-Eight Percent In. Sixteen Percent There.
Sixty-eight percent of manufacturers have started putting AI on the shop floor. Sixteen percent have hit the targets they set for it. Both numbers come from the same Boston Consulting Group (BCG) global survey of almost 1,800 manufacturing executives across seven industries (BCG, 2024).
If you have looked at that gap and concluded you picked the wrong tool, you are in good company. You are also, in a way the evidence now documents fairly clearly, wrong.
Most operations are working from a simple model: results track capability. Weak results mean the AI underperformed, and underperforming AI means switch vendors, widen the pilot, wait for the next model. That model is easy to believe. It is what a category built on capability releases talks about all day, and the tools really were thinner two years ago than they are now.
Two years on, the gap has not moved.
Grant Thornton went back and measured it. Their 2026 AI Impact Survey ran February through March across 950 business leaders, with a manufacturing subgroup of 100. Among those 100 manufacturers, zero reported a significant revenue increase from AI. Zero reported meaningful cost savings. Sixty-four percent reported efficiency gains (Grant Thornton, 2026).
Read those two numbers next to each other for a second. The AI was working. The money never showed up.
This is not something going wrong at your plant specifically. It is the shape of the whole category. The manufacturers who hit their targets in BCG's survey were not running better AI than the 84 percent who missed, and two years of genuinely better AI has not closed the distance between them.
So the useful question is not which AI to run. It is harder than that: what did the 16 percent actually do differently?
Why "Get Better AI" Feels Like the Right Answer
You have probably heard the standard advice, or given it. Expand the pilot to more lines. Move to a more capable model. Feed it more of your own data. Wait for the category to mature into something factory-grade.
None of that is bad-faith advice. It is just the only advice a capability-driven market knows how to give. Vendors publish capability benchmarks. Trade press covers capability releases. When you ask a peer at another manufacturer how their AI is going, the answer you get is which platform they bought. Capability is what everyone talks about because capability is the one variable the sellers control.
So if you deployed something, saw the cycle times improve, and then watched the improvement stop well short of the income statement, "the tool wasn't ready" is the most coherent explanation on offer. If you tried ChatGPT and got answers so generic you spent more time fixing them than you would have spent doing the work, waiting for a better model is a perfectly reasonable read. Tools that are visibly changing invite that diagnosis.
Here is where it breaks. Two years have passed since BCG ran that survey. The models got better. Compute got cheaper. More manufacturing-specific implementations came to market. If capability were the thing separating the 16 percent from the 84 percent, that stretch should have moved the number.
It didn't.
That does not mean a better tool is never worth buying. It means capability is not the variable that differs between the two groups, and the standard prescription has spent two years optimizing something that was never driving the outcome.
What Two Years of the Standard Prescription Produced
The clearest thing in the Grant Thornton data is the pairing. Same 100 manufacturers, same survey, two adjacent questions: 64 percent saw efficiency gains, zero percent saw revenue or cost movement worth calling significant (Grant Thornton, 2026).
That combination is a diagnosis, not a disappointment. Efficiency gains are what AI produces when it runs well against whatever target you gave it. P&L movement is what you get when the target itself was worth hitting. Those are different outcomes with different causes. Point AI at anything and you can get efficiency. You only get money out the other end if the thing you pointed it at was where the money was.
So the AI didn't fail. It succeeded at something that was never going to move the numbers.
Manufacturing is not special in this. MIT's Project NANDA reviewed more than 300 disclosed AI initiatives across enterprise populations for its 2025 State of AI in Business research and found roughly 95 percent of enterprise generative-AI pilots produced no measurable P&L return, against an estimated $30–40 billion in spend (MIT Media Lab, Project NANDA, 2025). McKinsey's November 2025 survey lands in the same place from a different angle: 39 percent of respondents attribute any operating-profit impact at all to AI, and most of that 39 percent put the contribution below 5 percent of earnings (McKinsey, 2025).
Factory floor, financial services, anywhere else. Same signature. AI deployed before anyone ranked what it should be working on.
The Tool Is Never the Variable
AI runs faster against whatever priorities you hand it. Rank those priorities by margin impact and it multiplies a clear signal. Let them come from whoever was loudest at the last all-hands and it multiplies that instead, just as efficiently. The tool has no way to tell the two apart. It amplifies. It does not diagnose.
Vera Nieuwland, Director of Business Consulting at Kaufman Rossin, puts the consequence plainly: investment without a foundation produces "more pilots, not more scale" (Nieuwland, 2026). She means foundation broadly — data infrastructure, organizational readiness. Our reading is that an unranked set of priorities is one of those missing foundations, and behaves the same way. AI running well against a murky priority gives you contained local wins that look fine in their own report and disappear on the P&L. Scale needs those wins to compound across the operation, and they only compound if the first target was the right one.
There is a sharper version of this for older plants. McElheran, Yang, Kroff, and Brynjolfsson found that among older manufacturing establishments, walking away from structured production-management practices accounts for roughly a third of AI-associated productivity losses (McElheran et al., 2025). That is a precise finding about a specific population, not a claim about every manufacturer. But if your production systems predate the cloud, the implication is easy to read: when the management structure underneath is informal and driven by escalation, the AI inherits that structure. It cannot fix a foundation it cannot see.
Three patterns tend to show up together when an initiative underperformed on priorities rather than on tooling.
The use case was picked because it was feasible. Quality inspection is the usual candidate. It is tractable, demonstrable, easy to benchmark against. It may also be the highest-margin process in your operation, or it may not, and feasibility does not tell you which. Choosing on feasibility means choosing without a financial ranking.
Nobody could connect the gains to a line someone owned. When the debrief talks about cycle time and throughput but never about cost or margin, the initiative probably did well against its internal metric and missed the one that mattered. That is a target-selection problem showing up as a measurement problem.
The scope came out of the meeting where the loudest complaint won. A single escalation or a vocal department head is an organizationally convenient priority. It is rarely a financially ranked one.
None of those are signs the tool failed. They are signs a step got skipped before the tool was ever chosen.
Call that step discovery. In plain operational terms: knowing which problem is worth working on before you commit a team or a tool to solving it. Not the most visible problem, not the most recently escalated one, not the most technically tractable one. The one that pays the most if you fix it. The AI cannot produce that ranking for you. It does not know what matters to your P&L until somebody tells it.
The manufacturers who hit their targets had that ranking before deployment started. Three separate research efforts describe what it looked like.
What the Manufacturers Who Hit Their Targets Did First
Three groups went looking for what separated the manufacturers who achieved their AI targets from the ones who didn't. Different populations, different methods, no coordination between them. They found the same thing.
Gartner surveyed 353 data, analytics and AI leaders between November and December 2025 and published the results in April 2026 (Gartner, 2026). The survey was cross-industry, so manufacturing was not its subject. Its central finding: organizations with successful AI initiatives invest up to four times more, as a share of revenue, in foundational areas than organizations with poor AI outcomes (Gartner, 2026). The foundations Gartner names are specific. Data quality. Governance. AI-ready people. Change management. Not platforms, not model upgrades, not algorithm selection. The same survey found only 39 percent of technology leaders confident their current AI investments will improve financial performance (Gartner, 2026). Take it as corroboration across industries rather than manufacturing proof, but the differential points at the same mechanism the manufacturing research finds on its own.
McElheran, Yang, Kroff, and Brynjolfsson worked from U.S. Census Bureau data covering 2017 through 2021, tracking real manufacturing establishments through the first wave of industrial AI. As reported in MIT Sloan Management Review, their finding is that AI's effect on manufacturing productivity depends on what it was built on top of (MIT Sloan Management Review, 2025). Firms already further along on digital and data-infrastructure maturity saw real productivity gains. Firms without that foundation, especially older establishments carrying legacy systems, saw productivity declines associated with AI (MIT Sloan Management Review, 2025). The paper frames this as a productivity J-curve: deploy before the foundation is ready and you don't just gain less, you go backwards for a while.
Their data also shows AI adoption in manufacturing tracking with prior digital maturity, using cloud infrastructure as the marker, rather than with on-premises IT spend (MIT Sloan Management Review, 2025). That is a correlation about context, not an instruction to go migrate your infrastructure. What it says is that the manufacturers who got productivity out of AI were already further along before they started. The foundation was the accumulated data: standardized production records, digitized process data, measurement systems that produce something a machine can read. Cloud adoption tends to travel with that. It does not create it.
The third stream is the most pointed, because it sits in the same report as the gap itself. Alongside the 68 and 16 percent figures, BCG sets out how it allocates the work of making AI pay off: 70 percent of the effort to people and business transformation, 20 percent to the data and technology backbone, 10 percent to the algorithms (BCG, 2024). That split is BCG's own prescription rather than an independent empirical result, and it should be read that way. It is still a striking thing to publish next to your own survey findings. The firm that documented the gap is saying, in the same breath, that nine-tenths of what closes it sits outside the AI.
Most manufacturer AI spending goes into that last 10 percent. Vendor evaluation, model selection, platform configuration. The 70 percent is the work of deciding what the AI should be aimed at, and getting the people around it to operate differently once you know.

Gartner from enterprise survey data. A Wharton team from Census records. BCG from its own analysis. Three methods, three populations, one finding: AI results depend on the foundation underneath.
The 16 percent were not running better AI. They had built that foundation first.
The Question to Ask Before the Next AI Decision
All three research streams converge on one question you can ask before the next AI spend goes out the door.
Was the problem this AI is being pointed at chosen by evidence, or by noise?
Not "which process is AI-ready." That is a technical test. Not "which complaint came up most in the last operations review." That is a volume test. The question is whether somebody ranked the candidates by what they are worth before anyone aimed a tool at one.
There is an easy way to hear the difference. Two descriptions of the same initiative:
"We deployed AI to quality inspection."
That describes a capability. It tells you nothing about whether quality inspection deserved the deployment. It could have been chosen because it was tractable, because the VP of Quality escalated hardest, or because a vendor demo made it look good. The sentence hides which.
"We chose quality inspection because it ranked highest by margin impact against the five processes we evaluated, above scheduling and above procurement."
That describes a decision. The deployment followed from a ranking. Whether the AI ends up succeeding or failing, the target selection is visible and defensible.
Five questions will tell you which sentence your own operation is about to produce:
- Which process is this AI going to target, and why did it rank above the next two by margin impact?
- Is there a single P&L line someone is accountable for improving as a result?
- Was this use case chosen because it was feasible, or because it paid best?
- If it succeeds on its internal metric, what changes on the income statement?
- Who ranked the candidates, and what data did they use?
A team that answers all five is working from a ranking. A team that stalls on the first is working from noise.
That ranking step, sitting between "here are the processes we have" and "here is where the AI goes," is what discovery means in practice. It is not a data governance program or an infrastructure migration. It is one question, asked before the money moves: if we could only improve one process this quarter, and we had to choose on margin rather than on volume of complaints, which one would it be?
Most manufacturers who came up short did ask that question. They asked it in the debrief, when someone needed an explanation. Asked then, the answer is a post-mortem. Asked first, it is a plan.
You know how the debrief goes otherwise. Someone asks what went wrong. Someone points at the tool. Someone suggests a different vendor. None of that touches the actual gap. The more useful account sounds like this: we deployed against a process that ranked high on visibility and low on financial return, and the step that would have caught that — ranking the candidates before aiming anything — was missing.
That is not an admission of failure. It moves the conversation off tool performance, where nobody can ever be sure, and onto priority selection, where the problem has a method and the method has a fix.
Your problem isn't execution. It's discovery. AI won't fix a foundation problem. It amplifies it.
You can run that ranking yourself. Plenty of operations do, once, in a good quarter with a motivated team. Repeating it is what breaks. The next cycle comes around, the process has to be rebuilt from scratch, and there is less time than there was last time. Hephanos is that discovery layer made repeatable. You bring the business knowledge, it scores and sequences the candidates, and you get the same ranked account of what to work on first every quarter. It is not another tool for the stack, and the answer it produces is never "buy more AI."
One practical note for the conversation ahead. At a 50–200 person manufacturer, the next AI spend almost certainly runs through a CFO expecting a return in the low hundreds of percent and a payback inside eighteen months. Those expectations hold no matter which platform is on the table. An initiative that cannot say which ranked priority it is attacking, and why that one ranked highest, will not clear them however capable the AI turns out to be.
Rank what the AI is pointed at before you point it at anything. That ranking is the foundation investment that decides whether the next AI dollar multiplies a clear signal or an unclear one.
References
BCG (2024). Knapp, D., Phillips, D., Sullivan, S., Yurek, D., Doerbandt, S., Shubhank, S., & Hartwick, C. Shaking Up the Factory Floor with Digital and AI. Boston Consulting Group, June 13, 2024. Global survey of almost 1,800 manufacturing executives across seven industries; also the source of the 70/20/10 implementation split. https://www.bcg.com/publications/2024/shaking-up-the-factory-floor-with-digital-and-ai
Gartner (2026). Organizations with Successful AI Initiatives Invest Up to Four Times More in Data and Analytics Foundations. Gartner press release, April 16, 2026. Global survey of 353 data-and-analytics and AI leaders, November–December 2025; commentary attributed to Rita Sallam, Distinguished VP Analyst. https://www.gartner.com/en/newsroom/press-releases/2026-04-16-gartner-says-organizations-with-successful-ai-initiatives-invest-up-to-four-times-more-in-data-and-analytics-foundations
Grant Thornton (2026). Manufacturing Insights: 2026 AI Impact Survey Report. Grant Thornton, April 2026. Fielded February–March 2026 across 950 business leaders, including a manufacturing subgroup of 100. https://www.grantthornton.com/insights/survey-reports/manufacturing/2026/manufacturing-insights-2026-ai-impact-survey-report
McElheran, K., Yang, M., Kroff, S., & Brynjolfsson, E. (2025). The Rise of Industrial AI in America: Microfoundations of the Productivity J-Curve(s). Working paper using U.S. Census Bureau data, 2017–2021. https://mackinstitute.wharton.upenn.edu/wp-content/uploads/2025/04/McElheran-et-al.-Industrial_AI_April-20-2025.pdf
McKinsey & Company (2025). The State of AI in 2025: Agents, Innovation, and Transformation. QuantumBlack, McKinsey & Company, November 5, 2025. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
MIT Media Lab, Project NANDA (2025). The GenAI Divide: State of AI in Business 2025. July 2025. Systematic review of more than 300 disclosed AI initiatives, plus structured interviews and survey responses.
MIT Sloan Management Review (2025). Burnham, K. The "Productivity Paradox" of AI Adoption in Manufacturing Firms. July 9, 2025. Coverage of McElheran et al. (2025). https://mitsloan.mit.edu/ideas-made-to-matter/productivity-paradox-ai-adoption-manufacturing-firms
Nieuwland, V. (2026). AI Arrived on the Factory Floor Before the Foundation Did. Automation World op-ed, July 8, 2026. Author is Director of Business Consulting at Kaufman Rossin. https://www.automationworld.com/factory/digital-transformation/news/55389491/op-ed-ai-arrived-on-the-factory-floor-before-the-foundation-did