Jason Forston
Perspectives · File 02

The Toolbar Fallacy

Enterprises put tens of billions into generative AI and most of it moved no line on any P&L. The reason is not the model. It is that nobody changed the work.

Jason Forston//Artificial Intelligence & Operations// 3 August 2026//11 min read

Somewhere between thirty and forty billion dollars of enterprise money went into generative AI, and according to MIT’s NANDA initiative, roughly ninety-five percent of it produced no measurable effect on profit and loss. Not a small effect. No measurable effect. The report calls the split the GenAI Divide, and the most useful thing about it is not the headline number, which will be argued about for years. It is the reason the researchers give for the failure, because that reason is not technological at all.

Before I use that statistic, I want to be honest about it, because I think how a leader handles a convenient number tells you more than the number does. The MIT finding is preliminary and has not been peer-reviewed. It rests on roughly three hundred publicly disclosed initiatives, fifty-two organizational interviews, and around a hundred and fifty senior-leader surveys.1 That is a serious sample for qualitative work and a thin one for a universal claim, and several large-sample industry studies over the same period report considerably more positive outcomes. So the honest reading is not “AI does not work.” The honest reading is that there is an enormous and consistent gap between companies that deploy AI and companies that get anything from it, and the gap tracks something other than model quality.

What it tracks is whether the company changed the work.

Section 01The fallacy, named

The dominant failure mode I see is that organizations buy a better toolbar for a process they have no intention of altering. The tool is genuinely excellent. It is placed on top of a workflow designed in 2014 for humans with a different set of constraints. Everyone is trained. Adoption is measured. Adoption is even achieved. And the P&L does not move, because the tool made an unchanged step faster, and that step was never the constraint.

I call this the toolbar fallacy, and it is the same error the industry has made with every general-purpose technology it has ever adopted. In 1987, Robert Solow made the observation that defined a decade of economics: you could see the computer age everywhere except in the productivity statistics.2 The computers were real. The investment was real. The output gains took another ten years to appear.

The best explanation for why came from Paul David, who went back and looked at what actually happened when American factories electrified.3 Factories that simply pulled out the steam engine and dropped in one large electric motor got almost nothing. The reason is structural and worth sitting with. A steam-powered factory was built around a central power source, which meant every machine had to be positioned relative to a system of overhead shafts and leather belts, which meant the entire floor plan, the material flow, and the building itself were dictated by where the power came from. Swapping the engine changed the power source and left the architecture intact.

The gains arrived only when someone realized that electricity permitted a unit drive: a small motor on each machine. That single fact meant the floor could be laid out according to the sequence of the work rather than the geometry of a driveshaft. Factories were rebuilt, single-story and wide instead of multi-story and cramped, organized around material flow. That is where the productivity came from. It took roughly forty years, and it was never about the motor.

They did not get a better factory by buying a better engine. They got a better factory by admitting the building was the problem.

Erik Brynjolfsson and his collaborators put a number on the modern version of this. The organizational and process investment required to make an information technology actually productive tends to run roughly an order of magnitude larger than the spend on the technology itself, and because that investment is intangible and slow, measured productivity dips before it rises. They call it the productivity J-curve.4 Companies in the trough of that curve look like they are failing. Some of them are. The ones that are not are the ones doing the unglamorous work of redesigning how decisions get made.

Section 02The shadow economy nobody wants to discuss

Buried in the same MIT report is a detail I find more damning than the ninety-five percent figure. In over ninety percent of the companies surveyed, employees reported regularly using personal AI tools for work, on their own, in parallel with the official rollout that was producing nothing.1

Read that carefully, because it eliminates every comfortable explanation. It was not resistance to change. It was not a training gap. It was not a culture that fears technology. The workforce had already adopted the technology enthusiastically, at their own initiative, and had voted with their browser tabs. What the organization could not do was fold that behavior into how the work officially gets done.

If your people are getting more value from a consumer chat interface they pay for themselves than from the platform you spent seven figures deploying, the problem is not your people and it is not the model. It is that your sanctioned tool was placed inside a workflow, and their unsanctioned one was placed inside their actual job. Those are different objects.

Section 03What “redesign the work” actually means

At ReadyNet I have been explicit that we redesigned the operating model around AI rather than simply adopting the tools, and I get asked what that sentence means in practice often enough that it is worth spelling out. It is not a philosophy. It is a sequence, and it starts somewhere counterintuitive.

  1. Start from the artifact, not the tool. List every recurring output your organization produces on a schedule: the reports, the decks, the summaries, the updates, the trackers. For each one, write down who consumes it, what decision it changes, and how long it takes to make. Most teams discover that a meaningful share of their recurring output changes no decision at all. It is produced because it has always been produced. You do not need AI for those. You need to stop making them. Automating an artifact nobody acts on just makes waste cheaper, which makes it permanent.
  2. Attack the assembly, not the authorship. The bottleneck in knowledge work is almost never writing. It is gathering: finding the current numbers, reconciling two systems that disagree, locating the version of the document that is actually current, remembering what was decided in March. Tools sold on drafting speed are solving the visible ten percent. Point the capability at the assembly step and the cycle time on a weekly decision drops from days to hours, which is the only kind of change that reaches the P&L.
  3. Put the model where the judgment is not. This is the discipline that keeps the whole thing honest. Never at the decision; always at the preparation for the decision. A model that assembles a complete, sourced, contradiction-flagged briefing for a pricing call is worth an enormous amount. A model that sets the price is a liability with excellent grammar. The value of a good executive is compressed almost entirely into judgment under uncertainty, and that is precisely the part these systems are worst at and most confident about.
  4. Change the definition of done, or you have not changed anything. If the review step takes as long as the creation step used to, you have moved the work rather than removed it, and you have added a review burden onto someone more expensive. This is where most pilots quietly die. It requires deciding, in advance and in writing, what quality bar a given output has to clear and who is allowed to accept it. That is an organizational decision, not a technical one, which is exactly why it gets skipped.
  5. Measure at the P&L line, not at the seat. “Hours saved” is not a result; it is an anecdote with a number attached, and it never survives contact with an actual finance review. Measure the things you already report: cost per order, days from concept to launch, gross margin points, pipeline conversion, service calls per thousand devices. If you cannot connect the deployment to a line that already exists on your P&L, you have not built a business case. You have built an enthusiasm.
A worked example

The clearest result I can point to is not a chatbot. At ReadyNet we designed a product-led growth motion for the SaaS platform around a free tier covering a customer’s first twenty managed devices, which strips the risk out of evaluation. By the time a customer outgrows that tier, they already depend on the reporting and management features that cut their service calls and truck rolls, so monetization scales with their own growth through a fixed monthly fee, a per-device charge, and tiered feature gating.

The AI work sat underneath that: content production, localization, personalization, customer-support strategy, and campaign automation, running across ChatGPT, Claude, Grok, Gemini, Adobe’s tooling, Zapier, and Make. What it bought was not a headcount reduction. It was the ability to run a four-product launch cadence, a channel program across thirty-plus partners, and a full rebrand with the marketing organization exactly the size it already was. Output and quality both rose. Headcount stayed flat. That is the real shape of the return, and it is worth being precise about it because the popular version of this story is a layoff, and that is not what happened.

Section 04The headcount question, answered directly

I want to take this one head-on, because executives ask it privately and answer it publicly in whatever way is least costly, and the gap between those two answers is doing real damage to trust inside companies right now.

Here is what I have actually observed. In a company where the constraint is labor supply, AI reduces headcount. In a company where the constraint is throughput on decisions, which is most companies above about thirty people, AI does not reduce headcount. It reduces the amount of a given person’s week that is spent assembling inputs so that someone else can decide something.

That has a second-order effect that almost nobody plans for. When you strip out the assembly work, what is left is judgment, and judgment is unevenly distributed in a way that assembly work was not. A team of ten in which four people were carrying the analytical load and six were carrying the production load does not become a team of six. It becomes a team where the composition of the job has changed underneath everyone, and if you do not say that out loud and retrain deliberately, you will lose the wrong people. They will not leave because they were replaced. They will leave because the job they were good at stopped existing and nobody told them what the new one was.

Section 05Governance without paralysis

The other failure pattern is the opposite of the toolbar fallacy: an organization so worried about risk that it writes a forty-page policy, routes every use case through a committee, and ships nothing for a year while its employees continue using consumer tools on personal accounts with zero oversight. That is not caution. It is the appearance of caution purchased at the price of actual control.

In practice, two questions filter almost everything that matters:

  • Does this touch customer, employee, or regulated data? If yes, it goes through the same review your existing data handling goes through. You already have that process. Do not build a second one with the word AI in the title.
  • Does anything leave the building without a human accountable for it? If yes, name that human and write down what they are certifying. If nobody will put their name on it, it is not ready, and that has always been true of everything a company publishes.

Everything else is experimentation, and experimentation should be cheap, fast, and largely unsupervised. The alternative is not safety. The alternative is a shadow economy you cannot see.

Section 06What the five percent are doing

Aditya Challapally, who led the MIT work, gave a description of the successful minority that is almost aggressively unglamorous: they pick one pain point, execute well, and partner smartly.1 There is no secret model, no proprietary architecture, no unusual budget. Gartner’s complementary finding points the same direction from the other end, projecting that a large share of AI projects will be abandoned specifically for want of AI-ready data, which is another way of saying the failure happened long before anyone chose a vendor.5

Artificial intelligence does not make a bad process good. It makes a bad process fast, which is considerably worse, because now you get to be wrong at scale and on schedule.

I think the discomfort underneath all of this is that redesigning work is a leadership act, not a procurement act. Procurement is fast, legible, and defensible in a board meeting. Redesign requires admitting that the way your company currently does something is worse than it needs to be, that the people who designed it are still in the building, and that changing it will make several capable adults temporarily bad at their jobs. Buying a platform lets you skip all of that. It simply also lets you skip the return.

Twenty years in, across lending technology, enterprise software, global manufacturing, consumer wellness, and next-generation telecom, I have not once seen a technology deliver a result that the organization had not already decided to reorganize itself around. The electric motor did not rebuild the factory. Somebody did, and only after they stopped asking what the motor could do for the building they already had.

Jason Forston
Chief Marketing & Operating Officer · Provo, Utah

Notes & sources

  1. MIT Media Lab, Project NANDA, The GenAI Divide: State of AI in Business 2025 (preliminary, July 2025). Based on roughly 300 publicly disclosed AI initiatives, 52 organizational interviews, and surveys of approximately 150 senior leaders. The report is preliminary and has not been peer-reviewed; several larger-sample industry studies over the same period report materially more positive outcomes. Lead-author commentary reported by Fortune, August 2025. Source ↗
  2. Robert M. Solow, review of Stephen S. Cohen and John Zysman, Manufacturing Matters, The New York Review of Books, 12 July 1987. The origin of what became known as the Solow productivity paradox.
  3. Paul A. David, “The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox,” American Economic Review 80, no. 2 (1990): 355–361. The account of unit-drive electrification and factory reorganization draws on David’s work and on Warren Devine’s history of electrification in American manufacturing.
  4. Erik Brynjolfsson, Daniel Rock, and Chad Syverson, “The Productivity J-Curve: How Intangibles Complement General Purpose Technologies,” American Economic Journal: Macroeconomics 13, no. 1 (2021). The order-of-magnitude figure for complementary organizational investment relative to technology spend comes from Brynjolfsson and Lorin Hitt’s earlier work on organizational capital.
  5. Gartner has forecast that a majority of AI projects lacking AI-ready data will be abandoned through 2026, locating the failure point upstream of vendor selection.
Figures cited from the author’s own operating record are drawn from Texas Armoring Corporation (2008–2020) and ReadyNet Solutions (2020–present). Third-party research is cited as published; where a finding is preliminary, contested, or subject to selection effects, that is stated in the text rather than in a footnote.
Jason Forston
About the author
Jason Forston

Jason Forston is a marketing and operations executive who has held the CMO and COO mandates at the same company, at the same time, twice. At Texas Armoring Corporation he served as acting chief executive of a 60+ person global manufacturer, rejected the armoring industry’s fear-based category premise, and grew the business from $2M to $25M. At ReadyNet Solutions he led a full manufacturing reshore from near-zero to roughly 100% US production through COVID disruption and tariff exposure while launching four new product lines and rebuilding the commercial engine. Harvard master’s in technology and digital media design; Duke Fuqua MBA; BYU political science. He writes about the seam between the demand engine and the operating engine, which is where he has spent twenty years.

Also in this series