Here are two numbers that explain the whole AI market right now. 88% of organisations report using AI in at least one business function. Only 39% can attribute any level of EBIT impact to it, and roughly 6% qualify as high performers with 5% or more EBIT impact. That’s McKinsey’s State of AI, published November 2025, from 1993 respondents across 105 countries.
So almost everyone has AI, but almost nobody has the money yet.
The gap isn’t weak technology – it’s which technology companies picked. The demos that sell boardrooms are conversational and general. The systems that show up in a P&L are narrow, boring machine learning: a model trained on your own data to make one specific decision, thousands of times a day, at a cost per decision close to zero.
Scaling faster: growth, not just efficiency
The best firm-level evidence here is a Journal of Financial Economics study by Babina, Fedyk, He and Hodson (2024). Tracking AI investment across roughly 1000 US public firms, they found a one-standard-deviation increase in AI investment came with 15 to 20% additional sales growth over eight years, plus 21,9% employment growth and 22,4% higher market value.
And here’s the part almost nobody quotes: they found zero effect on sales per worker and total factor productivity.
Read that twice, because it flips the standard pitch. ML didn’t make these firms leaner – it made them bigger. They used ML to launch more products, serve more segments, and take share, the way a restaurant that finally standardises its recipes can open a second location instead of just cooking faster in the first one.
PwC’s 2026 Global AI Jobs Barometer, built on more than a billion job ads across 27 countries, points the same way: headcount grew 52% at the most AI-exposed companies versus 36% at the least exposed.

Precision you can actually measure
Forget benchmarks for a second. The cleanest proof that ML beats the tools most companies still use came from the M5 forecasting competition: 5507 teams forecasting 42 840 real Walmart daily sales series, judged blind.
The winning entries beat the best statistical benchmark (exponential smoothing with bottom-up aggregation) by 20 to 22,4%. Every one of the top 50 submissions beat it by more than 14%. Gradient-boosted trees, specifically LightGBM, were used by four of the top five teams.
That’s a fifth of your forecast error, gone, using a technique that runs on a laptop and has been production-grade for a decade.
The same pattern shows up in knowledge work. Across 5172 customer support agents (Quarterly Journal of Economics, 2025), AI assistance raised issues resolved per hour by 15%, with the biggest gains going to the least experienced staff. Across 4867 developers at Microsoft, Accenture and a Fortune 100 firm (Management Science, 2025), completed pull requests rose 26%.
Saving money on mistakes
This is the underrated one, because errors are a cost most companies have simply stopped noticing.
US manufacturers paid $30,37 billion in warranty claims in 2025, or 1,30% of product sales revenue, calculated from the SEC filings of more than 1,400 companies. In healthcare, a JAMIA study of 6930 paired glucose measurements across 60 clinics found a 3,7% manual transcription error rate. Not a vendor estimate. A measurement.
Every one of those errors is a decision a model could have checked.
What that looks like when someone does it:
- John Deere See & Spray. Vision models distinguish weed from crop at 15 mph and spray only the weed. 2025 season: over 5 million acres, roughly 50% less non-residual herbicide, about 31 million gallons of mix saved.
- Frito-Lay. Reused an existing vision system to estimate potato weight, avoiding roughly $300 000 in capital per production line across 35 US lines, plus over $1M a year from peeling optimisation.
- Siemens Senseye at BlueScope. Predictive maintenance on steel production saved roughly 2000 hours of unplanned downtime over three years and prevented 53 full process interruptions.

Computer vision: the most mature ML you’re probably not using
If you want the shortest path from ML to money, it’s usually a camera.
On MVTec AD, the standard industrial defect benchmark, PatchCore reached 99,6% image-level AUROC, and it was trained on non-defective parts only. That matters enormously in practice, because most factories can’t hand you 5000 labelled photos of a defect they’re trying to eliminate. Show the model what “good” looks like, and it flags everything else.
Compare that to the human baseline. A 2024 survey of 196 papers in Applied System Innovation cites human inspector error rates of 20 to 30% on complex visual inspection tasks. Treat that as directional, since it’s a secondary citation to older human-factors work. And it isn’t because inspectors are bad. It’s because attention is a consumable and a shift is eight hours long.
At scale: BMW’s Regensburg plant runs automated paint-surface inspection and robotic correction on up to 1000 vehicles a day. Foxconn’s live SOP verification hit 96% recall and 99% precision, lifting first-pass yield by 3%. The computer vision market sits around $23,6bn in 2025, forecast to hit $101,5bn by 2033 at 20,1% CAGR (Grand View Research), with inspection and quality assurance the biggest slice.

What engineers say, and why you should listen
Stack Overflow’s 2025 Developer Survey (49 019 respondents) is blunt. 84% use or plan to use AI tools, but favourable sentiment fell from 72% to 60% in a year, 46% actively distrust output accuracy, and the single biggest frustration, cited by 66%, is “AI solutions that are almost right, but not quite.”
Practitioners are equally unsentimental about where value comes from. A recurring line across r/datascience: after ten years of trying, nothing has beaten XGBoost and random forests for fraud prevention. Or, as one commenter put it, intelligence is being able to build deep learning models; wisdom is knowing you won’t need them for most business problems.
And the failures are real. RAND cites a figure of more than 80% of AI projects failing, roughly double the rate of non-AI IT projects, with root causes that are almost never technical: leadership misaligned on the problem, thin data, technology-first framing, no budget for deployment. An engineer on r/computervision summed up a successful factory build in one line: the model was the least of our problems. Lighting was.
We’d rather tell you that up front than sell you a pilot. ML projects don’t fail at the algorithm. They fail because nobody owned the data pipeline, defined the KPI, or budgeted for the boring 80%.
Where the humans go
ML doesn’t mainly replace people, it reassigns them, and most companies fumble that part. BCG’s June 2026 study of about 12 000 workers found 42% of regular AI users save 8 hours a week, but 66% get limited or no guidance on what to do with that time, and over half aren’t reinvesting it in anything strategic.
The most useful rule we’ve seen came from a small business owner, not a consultancy: rank automation candidates by boredom, not complexity. Invoice matching, ticket routing, document extraction, visual pass/fail, restocking forecasts. High volume, clear rules, errors easy to catch. That’s where machines earn their keep, and where your best people are currently being wasted.

Start narrow, prove it, then scale
One decision → high volume -> measurable error rate today. That’s how we build ML at Softeta: as an engineering problem with a KPI attached, not a technology purchase. If you have a process that’s high-volume, error-prone and expensive, bring it to a 30-minute call and we’ll tell you honestly whether ML is the right tool. Sometimes the answer is a rules engine and better data plumbing. We’ll say so.