Back to Syllabi
Week 2 Corporate ExperiencesAI StrategyLeadership

How Do We Know AI Is Making Our Job Better

7 min read

Get posts like this in your inbox weekly

Subscribe Free

The Banker Who Never Went Home

Before computers, an investment banker worked brutal hours building financial models by hand. Then the spreadsheet arrived. Lotus, then Excel. The promise was obvious: the machine does the arithmetic, the banker goes home earlier.

The banker did not go home earlier. Benedict Evans makes this point well, both in his widely shared writing on Excel and in his May 2026 conversation on Lenny's Podcast. The spreadsheet did not change the hours. The banker still works 80 hours a week. What changed is that where you once built one model, now you build twelve scenarios. We took a better tool and used it to produce more at the same grind, with no clear sign the work got any better. Just more of it.

That is the trap every company is walking into with AI. The two questions on the table are almost always "how do we produce more" and "how do we plug AI in," and we know how that movie ends. More output at the same cost is not progress. It is a higher bar and an unchanged life. The better question, the one the banker never got to ask, is whether the tool can reshape what the role actually does, so the work itself gets better.

When the Most AI-Forward Company Can't Find the Value

Start with your own experience. Has Uber gotten noticeably better for you this year? The app, the ride, the price, any of it? For most of us the honest answer is no, it feels about the same.

Now sit with this. Uber ran an internal leaderboard ranking teams by how much AI tooling they used, and burned through its entire 2026 AI coding budget in four months. The incentive worked perfectly. Usage went straight up. Enormous AI spend, every team racing up the chart, and the product in your pocket feels exactly the same.

Their own COO said as much. On the Rapid Response podcast this spring, Andrew Macdonald admitted the company could not draw the line. More was probably shipping, he said, but it was very hard to connect any of the usage numbers to actually producing 25 percent more useful consumer features. Microsoft, around the same time, quietly cancelled most of its Claude Code licenses in one division and pushed engineers back to a cheaper tool.

Here is the part everyone misses. Uber asked "are we producing more" and could not answer it. But producing more was the wrong question to begin with. The right one: did the work change, or did a lot of people just get faster at the old job?

Why the Easy Metrics Lie

When we want to know if AI is helping, we reach for what is easy to collect: how many people use it, how many features we shipped, how many tokens we burned. Every one of those is an activity count, and every activity count can be gamed. Put a leaderboard on usage and usage goes up. It tells you the tool is being touched. It tells you nothing about whether the work got better, or whether the role is better to do.

Anyone who has built a regression knows the trap. Throw more variables into the model and R-squared never goes down, it almost always ticks up, and the fit looks better on paper. But the model did not get smarter. You just gave it more to chase. That is why we use adjusted R-squared, which penalizes the padding and asks whether each addition actually earned its place. Usage dashboards are R-squared with no adjustment. The number climbs because you added inputs, not because the output got better.

So what is the adjusted version?

The Process Was Built Around a Constraint That No Longer Exists

Take how we used to validate a product idea. Someone pitches it, we hold a kickoff, we spend two weeks on a Figma board nobody can click through, we collect feedback on the pictures, then we spend another week rebuilding it as something interactive. Four weeks in, the excitement has cooled and three new priorities have surfaced. The idea limps to a decision, if it gets one at all.

The fix was not "use AI to build the Figma board faster." That is the banker's bet again, a faster tool bolted onto the same broken sequence. The detailed task list was never the problem to optimize. It was the problem itself. Disney's Imagineers do not start with a task list, they start with the experience they want and engineer backward to produce it, the way Walt treated Disneyland itself as a prototype he walked and reshaped in real time. The point is to stand inside the working thing instead of reviewing a drawing of it.

That reframed our whole sequence. Every step was scaffolding built around one constraint: a working, clickable prototype used to be expensive. The kickoff, the static mockups, the staged feedback, all of it existed because you could not put the real thing in front of people on day one. Once a working prototype became cheap, that constraint was gone, and the process built around it no longer made sense. So we threw it out. Now the first meeting is the prototype. We walk in with a working version, and that meeting becomes a two-hour workshop instead of a kickoff. We talk through it live, elements get changed in the room, and we leave with real alignment and a focused direction for what to build next.

The Answer Isn't Black and White

We probably cannot yet prove AI made the product better, and anyone who claims they can is usually pointing at a usage chart. Even the tools themselves make this hard to judge. Claude ships updates almost daily, and on any given day it is genuinely tough to tell whether the output got better or just got faster at producing slop. The signal is noisy, and pretending otherwise is how you end up with a leaderboard and a burned budget.

What we can see this early is whether the work changed shape and whether the people doing it operate differently than they did a year ago. We are moving into a world of workflows, not tasks. The whole reason usage and token counts mislead is that they measure tasks, how many times the tool was invoked, in a world that is quietly reorganizing around workflows, how the work actually flows from idea to outcome. Count tasks and you will always miss the thing that moved. That is not a victory lap, and it is not the same as "people feel more engaged," which is just usage wearing a friendlier face. It is the one honest signal we have that the work itself changed, not the speed at which we do it.

So stop asking how to produce more and how to plug AI in. The banker produced more for forty years and still never went home. The better question is not how to build faster, but what the work should produce, and how to reshape it to get there. AI hands you the banker's path by default. The whole job now is refusing it.


Sources

Benedict Evans, "A rational conversation on where AI is actually going," Lenny's Podcast, May 31, 2026. https://www.lennysnewsletter.com/p/a-rational-conversation-on-where

Benedict Evans, "AI and the automation of work" (on Excel, the Jevons paradox, and accounting employment), ben-evans.com, July 2, 2023. https://www.ben-evans.com/benedictevans/2023/7/2/working-with-ai

Jason Del Rey / Fortune, "Uber burned through its entire 2026 AI budget in four months. Now its COO is questioning whether it's worth it," May 26, 2026. https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/

Andrew Macdonald interview, Rapid Response podcast (Uber COO on AI spend and consumer features), May 2026.

The Street, "Amazon joins Microsoft in sending shocking message to employees" (Microsoft Claude Code license cancellations; Amazon and Meta leaderboard shutdowns), May 31, 2026. https://www.thestreet.com/technology/amazon-joins-microsoft-in-sending-shocking-message-to-employees

Justin Grosz

Justin Grosz

Product Leader | Adjunct Professor, Northeastern