AI 20/20 Subscribe

FIELD NOTES · ENTERPRISE AI

Enterprise AI: what shipped

Most enterprise AI announcements never reached production. This page tracks the record instead of the announcements: named companies, linked public sources, what each deployment did, what it took, and what broke. The failures are on the list on purpose.

COMPILED 2026-08 · EVERY CASE LINKED TO A PUBLIC SOURCE · UNSOURCED CASES CUT

Your skepticism about enterprise AI is earned, and the data agrees with you. MIT's 2025 State of AI in Business review, built on 300 public deployments and 52 executive interviews, found that 95 percent of enterprise generative AI pilots produced no measurable P&L impact, against an estimated $30 to $40 billion of investment.

Which makes the remaining 5 percent the only part worth studying. Below are deployments that actually reached production at named companies, each linked to a public source. Two rules govern the list: a case with no public source gets cut, not paraphrased, and where a number comes from the company or its vendor, the page says so. Reaching production is also not the end of the story, so the cases that shipped and then broke stay on the record with their endings attached.

The record

5 CASES

CASE 01 · WEALTH MANAGEMENT · SHIPPED

Morgan Stanley: the advisor assistant

What it did

A GPT-4 assistant that answers advisor questions against the firm's own research library, rolled out to roughly 16,000 advisors from 2023. A second tool, Debrief, sits in client meetings with consent and drafts the notes, action items, and follow-up emails into the CRM.

What it took

An evaluation suite that gated every rollout, built with the model vendor before launch. Per the vendor's case study, retrieval coverage of the document corpus went from 20 to 80 percent, and adoption reached 98 percent of advisor teams. Those figures are the company's and the vendor's own. Treat them as their claim.

What broke

Nothing on the public record so far. The design choice worth copying: the assistant retrieves and drafts, the advisor sends. The human stayed on every output that touches a client.

CASE 02 · BANKING · SHIPPED

JPMorgan: LLM Suite

What it did

An in-house portal giving employees access to large language models inside the bank's own security and compliance controls, for drafting, research, and analysis. From zero to 200,000 onboarded employees in eight months across 2024 and 2025.

What it took

Building where the data controls already lived, and treating adoption as something to earn: the rollout was optional first and woven into existing workflows rather than mandated. It also took a bank-sized technology organization, which is the honest footnote. MIT's review found externally bought tools reach production about twice as often as internal builds. JPMorgan is what the exception costs.

What broke

Nothing on the public record so far. The bet is scale over specialization: one governed doorway to models instead of hundreds of scattered pilots.

CASE 03 · CUSTOMER SERVICE · SHIPPED, THEN CORRECTED

Klarna: the assistant and the sequel

What it did

An OpenAI-built customer service assistant that handled 2.3 million conversations in its first month of 2024, two thirds of all support chats, and cut resolution time from 11 minutes to under 2. The company projected a $40 million profit improvement for the year. Those are Klarna's own figures, from its own press release.

What broke

Quality, slowly. By May 2025 Klarna was rehiring human agents for disputes, complex refunds, and hardship cases, with the CEO conceding publicly that cost-driven automation had produced lower quality service. The model that survived is hybrid: the assistant handles the routine volume, a human takes over where the cost of an error rises.

The lesson

The launch press release and the correction are the same case study, fourteen months apart. Any vendor citing Klarna's month-one numbers without the 2025 sequel is telling you half a story, which is a useful thing to know about the vendor.

CASE 04 · QUICK SERVICE · ENDED

McDonald's: drive-thru voice ordering

What it did

Automated order taking, built with IBM and tested from 2021 at more than 100 U.S. restaurants. A real production deployment, at scale, facing customers in the open air.

What broke

Accuracy in uncontrolled conditions. Background noise, accents, and overlapping voices produced order errors that staff had to redo, and that customers filmed. The test ended in June 2024 and the technology came out of every restaurant by the end of July.

The lesson

Ambient, customer-facing audio is a different problem from a demo booth, and two years of production testing is what it took to establish that. The failure is more informative than most success stories: it puts a boundary on the map.

CASE 05 · CUSTOMER SERVICE · LIABILITY SET

Air Canada: the chatbot and the tribunal

What it did

A website chatbot answering policy questions. In 2022 it invented a bereavement refund policy that contradicted the airline's actual terms, and a customer booked on the strength of the answer.

What broke

The legal theory that the bot's words were not the company's. British Columbia's Civil Resolution Tribunal ruled in February 2024 that Air Canada was responsible for everything on its website, chatbot included, rejecting the argument that the bot was a separate entity. The award was small, CA$812. The precedent was not.

The lesson

Your AI's output is your company speaking. Every customer-facing deployment on this page now operates under that ruling, which is why the question of what happens when the system is wrong belongs in the first vendor meeting, not the post-mortem.

What the ones that shipped share

THE PATTERN
P.01

A narrow job. Answering advisor questions. Drafting meeting notes. The deployments that held pointed AI at one bounded task; the ones that broke pointed it at open-ended contact with the public.

P.02

A human kept approval where the cost of error rises. Morgan Stanley's assistant drafts and the advisor sends. Klarna's correction was re-inserting the human at exactly that line. The pattern is the same one drawn from both directions.

P.03

Measured against the old workflow, with evals before rollout. The shipped cases could say what the process cost before them. The 95 percent, mostly, could not.

P.04

Built where the data and the controls already were. Inside the research library, inside the bank's compliance perimeter. Not bolted onto the outside of either.

P.05

Adoption was earned, not mandated. Optional first, woven into the existing day, expanded on demand. The tools people route around do not appear on this page, because routing around a tool leaves no record.

P.06

A stated boundary. Every failure on this page happened at an edge nobody had named in advance: noise at a drive-thru window, a grieving customer with a policy question. The shipped cases knew their edges. The broken ones met theirs in public.

Where your company sits

SELF-LOCATION

Read your own AI effort against the record. If your pilot has no baseline measurement and no agreed metric, it is in the 95 percent already, whatever the demo looked like. If your deployment faces customers and nobody has written down what happens when it is wrong, Air Canada has already run that experiment for you. If the plan assumes the launch-month quality holds forever, Klarna's sequel is the base rate.

And if your effort has a narrow job, a measured baseline, a human on the consequential outputs, and a named boundary, you are working the same pattern as the cases that shipped. The remaining questions are vendor questions, and the list for those is on the vendor evaluation page.

THE NEXT STEP

What shipped, as it ships

A record like this is only useful while it's current. I send one email on what I built, what it cost, and what broke. The same work updates this page.

 

Keep reading

THE CLUSTER