AI Did Not Break Productivity. It Exposed Our Decision Problem.

James Eselgroth • June 16, 2026

BLUF: AI does not remove the need for judgment. It raises the cost of getting it wrong.

I have watched a strange pattern in leadership rooms. When the decision is small, everyone has an opinion. When the decision is large and consequential, the room often goes quiet.

I saw it in the Air Force. Bring up a handgun, and everybody leans in. Four-star generals. Colonels. Senior NCOs. Airmen. Everyone has fired one. Everyone has a view. The debate runs long because the object is familiar.

Now bring up a major weapon system. The physics get harder. The trade-offs sit buried under engineering, doctrine, logistics, and cost. Suddenly fewer people challenge the framing. The decision feels too complex to question.

So the room moves faster.

That is the dangerous part. Complexity should breed more discipline. Too often it breeds false confidence.

AI adoption exposed the same habit. Many organizations saw the promise clearly. AI could write code. AI could summarize documents. AI could automate work. AI could cut cost. Each claim held some truth. But truth without context becomes risk.

The motion trap

Noah Byteforge's recent piece on Microsoft, Uber, and Klarna captured the tension well. His sharpest point was not whether AI can produce output. It can. The problem was that leaders could not connect those output metrics to actual user value. Commits and usage statistics looked impressive. Customer outcomes were harder to prove.

That is not an AI problem. That is a decision problem.

Organizations measure what they can see. A commit is visible. A token count is visible. A usage dashboard is visible. Customer value is harder. System understanding is harder. Judgment is harder.

The old warning still holds. Do not confuse motion with productivity.

Klarna learned the difference the hard way. The company replaced roughly 700 support roles with an AI chatbot and let it handle most customer interactions. Then satisfaction dropped, and by Byteforge's account the firm began rehiring human agents. The AI handled the patterns it knew. The judgment cases piled up with no one to catch them.

Motion feels good. It creates evidence. It makes the board slide easier. It makes the transformation story sound active. But activity is not progress. Progress requires a link between effort and outcome. That link is where most AI efforts are weakest.

The engineer is not a code factory

An engineer is not valuable because they produce lines of code. That is the visible artifact. The deeper value sits underneath. They carry system memory. They remember why a shortcut failed three years ago. They know which dependency looks clean but causes pain. They hear the risk inside a casual meeting comment.

AI can read the code. It cannot read the meeting where a team rejected an approach after a failure that no one documented. It cannot recover a decision history that was never captured.

That missing history matters. It is part of the decision supply chain. Inputs come in. Assumptions get made. Trade-offs happen. Decisions become work. Work becomes systems. Systems shape the customer experience.

When decisions are not captured, the organization loses memory. When AI enters that environment, it does not fix the gap. It accelerates it.

The amplifier

AI makes strong systems faster. It makes weak systems louder.

If leaders measure the wrong thing, AI optimizes the wrong thing. Target commits, and teams produce commits. Target token usage, and teams burn tokens. Target automation volume, and people automate work without asking whether the work matters.

That is theater. It looks modern. It produces dashboards. It does not produce value by default.

The model predicts the word you want

Here is the trap underneath the trap. These models do not reason. They predict. And they do not predict the next word in general. They predict the next word you want.

The model tailors to its user. It builds a profile. The next word it offers the President is shaped differently than the next word it offers a high schooler. That personalization is useful. It is also a bias engine if you do not see it working.

Throw a question at the black box without context, and the model fills the gaps with assumptions. It will not stop to ask you about them. Each unexamined assumption stacks on the last. In business vernacular, a stack of unexamined assumptions is risk.

The accountability does not transfer to the tool. No organization tolerates “the AI said so” as an excuse. Use AI to help, and you still own the output. That is the part too many leaders skipped on the way to the demo.

A simpler tool was available

Here is the part that should sting. The tool to think this through already existed. Before turning the AI dial up or the headcount dial down, a leader could have mapped the decision.


A Causal Decision Diagram does this. It is not a flowchart. It shows the levers you control, the chain of effects they set off, and the measurable outcome you actually care about. A lever is a quantity you can move up or down. You start simple. One lever, one chain, one outcome.

Figure 1. First-draft CDD: one lever, the chain it runs through, the outcome it should reach.

That first draft looks reassuring. Turn AI usage up, output rises, more features ship, more value. One clean line. It is also incomplete, which is exactly how every real CDD starts.

Now add what the first draft hid. Keep both levers neutral. AI usage and engineering headcount are dials you can turn either way. Then follow the arrows.


Figure 2. Wired CDD: neutral levers, leading indicators, external, outcome, goal, and inverse dependencies marked.

Read it left to right. Turn AI usage up and code output climbs. More output means more code to review. But review coverage is the share of that code a human actually checks, and the team has only so many hands. As output runs ahead of the team, coverage drops. Thinner coverage lets more defects slip through. More escaped defects, lower customer satisfaction.

The second lever drives the same outcome from below. Engineering headcount feeds two things: review capacity and system memory. Cut headcount and both shrink. Less capacity thins coverage further. Less system memory means fewer people who remember why the code is shaped the way it is, the context that catches the defect no document explains. Both losses push defects through. Both pull satisfaction down.

So the diagram makes the real rule visible. Turn the AI dial up, and you have to turn headcount up too, or coverage collapses and the outcome turns. That is not an argument against AI. It is the cost of using it without paying for the review and the memory it leans on.

Now look at the box set apart at the top. Token utilization is what teams count. It branches off code output, easy to measure, easy to chart. Notice what is missing. No arrow runs from it to customer satisfaction. The most-watched number on the dashboard has no path to the thing that matters. That gap is the whole article, drawn in one diagram.

The diagram does not make the decision for you. Good. That is not its job. Its job is to slow the room down and trace the chain before you pull the lever. The measurement that matters sits on the intermediates, the leading indicators. Review coverage and defect escape warn you months before the satisfaction number moves.

That is the difference between measuring motion and measuring outcomes, drawn in one picture.

Augmented intelligence is the operating model

The future is not human versus machine. That framing is lazy. The future is human judgment paired with machine execution.

Chess settled this years ago. A machine beats a human. But a human-machine team beats a machine. The humans shape strategy. The machines sharpen execution. The value lives in the pairing.

The same holds now. People bring context, ethics, accountability, and the judgment to read the arrows. AI brings speed, scale, and generation. A leader using AI without understanding the business is not transforming anything. They are outsourcing assumptions and calling it progress.

The hard work remains. Understand the business. Understand the decision. Understand the system. Understand the outcome. Then use AI with purpose.

The companies that win with AI will not replace judgment with generation. They will preserve judgment and accelerate it.

That requires harder measurement. Not how much code was written. Not how many prompts were run. Measure whether customers were better served. Measure whether decisions improved. Measure whether the system got stronger.

AI did not break productivity. It revealed how poorly many organizations understood it.

Do not confuse tokens with productivity.