Same AI model, so why does one company get compounding returns and another gets nothing
- A salesperson gets a full client brief seconds before the phone rings, because AI is plugged into the company's real data streams. Those few seconds capture what this whole piece is about
- An Anthropic e-book names the gap between "using AI" and "getting compounding returns from AI": the agentic thinking divide. After adoption doubled in two years, who's using it no longer differentiates anyone
- Point solutions are inevitably mediocre: generic AI gives generic output, employees get something they "still have to fix," and the gap isn't about the model — it's about how much organizational context you feed in
- Each of the three pillars has real evidence: L'Oréal's conversational analytics hit 99.9% accuracy, Lyft cut support resolution time by 87%, Rakuten went from a major release every quarter to every two weeks
- This advantage compounds: expert feedback flows back into the knowledge base, the capability curve climbs; starting a year late doesn't just cost you a year — it costs you a year of compounding
- At the end, the four practical principles and the six-month three-phase timeline are quoted in full — you can copy them directly
What happens in the seconds before the phone rings
A salesperson sits at his desk. The phone is about to ring in a few seconds. On the other end of the line is a client he's been chasing for half a year.
Three months ago, he'd have been scrambling: five or six tabs open, digging through the CRM for this company's past interactions, hunting through meeting recordings for what was said last time, then switching to a research tool to check whether the client had raised funding recently, or whether a competitor had landed on their vendor list. This manual grind used to take hours, and he'd usually only manage a quick glance before the phone rang.
This time, he typed a command before the call. A few seconds later, a brief appeared in front of him: the company's latest data, his entire interaction history with this contact, exactly where the unclosed deal was stuck, and what competitors had recently been pitching them. He picked up his water glass, calm, and answered the call.
Enterprises are long past the "should we use AI" question
Anthropic published an e-book for enterprises called Building AI agents for the enterprise, subtitled "Best practices from industry leaders." It puts on the table something that's been quietly happening but rarely gets said out loud.
When a technology's adoption rate doubles in two years and is still accelerating, who's using it and how much no longer differentiates anyone. What actually separates companies from each other is a different question: are you treating AI as a tool sitting on a desk, or as the underlying capability that reorganizes your entire company.
This book pulls that variable out and gives it a name: the agentic thinking divide. Below, we follow its two hardest claims — why scope decides everything, and why this kind of advantage compounds — and unpack them layer by layer. You don't need to memorize the conclusions; by the end you'll be able to derive them yourself.
Adoption vs. transformation: the difference is obvious at a glance

In one line: point solutions only ever give you point results. Each one looks fine on its own, the demos can even be impressive, but they share a common fate: stuck forever at the pilot stage, changing nothing about how the company actually runs.
A chatbot is a question-answering machine; an agent is a coworker who gets things done. You ask a chatbot a question, it answers, then forgets. An agent is handed a goal, and it breaks it down, decides, and executes step by step until it's done — adjusting along the way based on results. A lot of so-called "AI transformation" is really just a row of question-answering machines lined up. You cannot build transformation out of chatbots — that's a first-principles limitation.

Why point solutions are inevitably mediocre
Let's state the claim up front: companies that treat AI as an isolated tool are destined for mediocre results. This isn't a matter of attitude — it's structural.
These tools run on "generic AI," and generic AI produces generic output: something grammatically correct, structurally complete, but that anyone looking at it thinks "this still needs work." That "still needs work" is exactly the problem. An employee gets a document drafted by AI and finds it doesn't understand the company's standards, doesn't use the company's terminology, doesn't know the institutional knowledge that lives only in veteran employees' heads and was never written down anywhere. So the employee still has to spend time refining it until it's usable.
This is the difference between "AI drafts a document" and "AI drafts a document your team can ship as-is." It sounds like a few words apart, but what sits in between is an entire company's worth of context. The two aren't a gap in model quality — they're a gap in how much context you fed it.

What's the evidence? A pattern shows up repeatedly in the book: two companies use the exact same model and get wildly different results — the difference is how much organizational context got fed in. That one sentence already kills "model determinism" outright: if the same model can produce completely different results, the model clearly isn't the deciding variable. The deciding variable is scope. Treat AI as a point tool and you only get point results; treat it as transformation and you're rebuilding three things at once: how employees work, how processes run, and what products you can build.
Those three things are the book's core framework: the three pillars. Let's take them one at a time — a real company stands behind each one.
L'Oréal: moving everyone's starting line forward
Back to that salesperson at the start. What made those few seconds possible wasn't a smarter model — it was the model being plugged into the company's real data streams: CRM, meeting recordings, prospect research. Swap in a different company with a different dataset, and the same model can't produce that brief.
This isn't unique to sales, either. Finance connects to the data warehouse to produce reconciliation reports; legal reviews contracts against the company's own risk framework and flags deviant clauses; marketing drafts campaign proposals against brand guidelines. The pattern holds: the value scales with how much organizational knowledge you've encoded into it.
L'Oréal put a number on this pattern. Picture its situation: products sold in over 150 countries, data scattered across countless pipelines. An employee wanting to know how a particular product line is selling in a particular market has to file a request, wait for a data specialist to build a custom query, and hope the specialist understood what they were actually asking. The whole company's data capability was bottlenecked right there.
Its solution was an internal AI platform built on Claude: a multi-agent system that routes plain-language employee questions automatically to the right data sources and 15-plus specialized agents, then synthesizes the results into an answer with charts.

Note what the key point here is not: it isn't "swapped in a stronger model." It's 90 approaches tried against a 99.9% that finally worked — the one that worked was the one that correctly orchestrated 15 specialized agents against the right data sources. Orchestration and context are what filled the gap between 90% and 99.9%. Thomas Menard, who leads L'Oréal's agentic platform and LAB, put it this way: "Our automated evaluation capability, like LLM-as-a-judge, has repeatedly proven the superiority of Claude models."
Lyft: compressing "months" into "minutes"
This pillar tells a different story: it's not one person getting faster, it's an entire pipeline getting faster. And there's a counterintuitive pattern: the more complex and information-dense the process, the bigger the payoff.
The reason is the same one as before: the value of process automation depends entirely on the context behind it. Build standards, compliance requirements, and institutional knowledge into the system, and processing time can drop from months to minutes without losing quality. What success looks like: clinical writers go from spending weeks stitching together a report to spending most of their hours reviewing and polishing; compliance officers go from spending days producing regulatory filings to generating a first draft in minutes; a whole team goes from spending 80% of its time producing documents to spending 80% of its time making judgment calls.
Capacity shift: AI isn't replacing people, it's moving people from "production" to "judgment." The machine takes over the manual labor, freeing people up to do the things only people can do.

Lyft is the hard evidence for this pillar. If you've ever run into trouble on a ride-hailing app late at night, you probably know the frustration: something goes wrong with your ride, you wait 30-40 minutes to get through to support, and the agent on the other end is juggling three or four people at once with copy-pasted template replies. That was exactly Lyft's situation: spanning six continents and thousands of cities, its support system pushed to the breaking point, with agent burnout climbing right along with it.
Lyft was thorough in choosing Claude: it tested both raw performance and whether the tone matched the brand voice. It started with driver support, then expanded to rider support and billing disputes. The setup now: Claude greets the customer by name, investigates the specific situation, and resolves it in seconds; only when a case genuinely needs human judgment does it get routed to a human agent, along with a summary of the conversation Claude generated itself.
Note the causal chain here: the 87% drop in resolution time isn't because Claude types faster than a human agent — it's because it's embedded in Lyft's entire ticketing workflow, able to investigate, judge, and hand off to a human with a summary when it should. This is the payoff of embedding depth, not the payoff of model speed.
The savings run into the millions of dollars, and Lyft didn't pocket them — it reinvested them into the support team and new initiatives, one of which is Lyft Silver, dedicated one-on-one support built for older riders. AI freed people from mechanical labor, and the savings turned into a warmer new service that simply couldn't have existed before.
Rakuten: letting customers do what they couldn't do before
The first two pillars looked inward; this one looks outward: AI doesn't just save you money — it lets you build products that simply couldn't have existed before.
The book points to a shared pattern: frontier AI models + proprietary data + existing trust relationships + deep domain expertise. AI is just the enabler; the real moat comes from everything around it. So the product-layer opportunity is never just "cut costs" — it's using new product capability to create net-new revenue and a compounding competitive edge: whoever moves first builds the integrations, the data flywheel, and the customer habits that make it hard for latecomers to catch up.
Trust boundary: the scope of data a customer is willing to hand over for you to process is the trust boundary. In regulated industries like finance and healthcare, data security and compliance aren't nice-to-haves — they're the price of entry. Any AI product that operates outside the trust boundary is, functionally, a product that never ships.

The company that builds out this pillar most fully is Rakuten. It runs more than seventy businesses and has pushed an "AI-nization" strategy company-wide. It understood early on that for agents to really get work done, they need persistent compute, memory, and storage. So its engineers initially built the infrastructure from scratch. That call was right at the time, but it came at a cost: top engineering talent that could have gone into differentiated innovation was instead spent entirely on laying the groundwork.
The turning point was adopting Claude Managed Agents (a pre-built, configurable agent runtime framework provided by the Claude platform), outsourcing the entire "execution layer" grunt work so its own engineers could go back to focusing on genuinely agentic experiences. The effect was like a floodgate opening: within a week, specialized agents covering engineering, product, sales, marketing, and finance were deployed, wired directly into Slack, Microsoft Teams, and the company's own dashboard systems, running long tasks that lasted hours. And the agents' memory compounds: they remember mistakes made in the past, so they don't repeat them.
But what makes Rakuten most memorable isn't the numbers — it's one image. Internally, they call their cross-domain power users "Galileos." One of them is a product manager, not an engineer. Alone, he built a FinOps (financial operations) pipeline across several public clouds, and set up his own monitoring agent running quietly in the background. This used to be something a product manager wouldn't have dared to attempt — they'd have had to line up and beg the engineering team for it. That's the structural conclusion Rakuten offers: an agent isn't your future coworker, nor is it a competitor coming for your job — it's infrastructure the company uses to accelerate building everything.
This advantage compounds — it doesn't grow linearly
The first two claims are already established: point solutions are mediocre, transformation runs on organizational context. But that's not enough. If the advantage were just a one-time "I'm a bit better than you," it would eventually get matched. The book's sharpest claim sits at this third layer: this advantage snowballs — the distance between first-movers and everyone else keeps growing.
Where does the compounding come from? The book gives a very concrete mechanism, "how accuracy compounds." The common approach: have experts review AI's output against the same baseline every time — the experts wear themselves out, and the AI never improves. The right approach: build a system that feeds human expert feedback back into the AI's knowledge base, so every expert review makes every future process a little better. That means the AI's capability curve climbs upward — it isn't flat.

It's like a veteran employee mentoring a new hire. The experience from one round of mentoring settles into a playbook, so the second new hire doesn't need to be taught from scratch — one person's learning instantly becomes the whole organization's learning. Rakuten's line earlier, "the agents' memory compounds," is the literal implementation of this mechanism: a hole one agent falls into is a hole no agent falls into again.
Separating "tool" from "infrastructure" is the key to understanding compounding: a tool is something you use and put back down, its value fixed; infrastructure is something you keep building on top of, its value rising the more you build on it. Point solutions are tools; transformation is laying infrastructure. That's why the former is linear and the latter compounds.
Follow this mechanism through and the conclusion is direct: the organizations that start earliest accumulate the biggest advantage. Because every month's worth of accumulated approval records, expert feedback, and correction cases makes next month's output faster and more accurate. Starting a year late doesn't mean falling behind by a year — it means falling behind by "a year's worth of compounding": for that whole year, your competitor got stronger every month, while you're starting from zero.
Plugins turn institutional knowledge into organizational infrastructure
The chain of logic is closed at this point: point solutions are mediocre → transformation runs on context → context compounds → so first-movers win. But this chain has an implicit premise: looking back at every case above, L'Oréal's 15-agent orchestration layer, Lyft's support system, Rakuten's Managed Agents, and also Novo Nordisk (mentioned in the book — NovoScribe compressed a clinical study document from 10-plus weeks to 10 minutes) and RBC (an agentic solution serving 2,200 advisors managing $689 billion in assets) — every one of them built a custom platform of their own, specifically to feed Claude context.
The payoff is substantial, but the barrier is real too: every one of them required engineering resources, time, and technical expertise. The starting line for this game was set at "you need an engineering team that can build a custom platform." The vast majority of knowledge workers were shut out.
Claude Cowork changes that equation. It gives non-technical people access to the same agent capability that enterprises built for themselves on the API — without any custom development. You hand it a task, and what it gives back isn't a suggestion or an outline — it's genuinely finished work: a Word document, an Excel model, a slide deck, an analysis report.
The mechanism behind it is plugins: skills, context, and connectors packaged into a single plugin, giving Claude a role-specific expertise. Institutional knowledge that used to be locked inside a veteran employee's head — gone when they left the company — is now packaged into a plugin: copyable, shareable, accumulable. It effectively engineers "individual learning becomes organizational learning." Anthropic has already open-sourced 11 plugins, covering productivity, sales, finance, data, legal, marketing, customer support, product management, enterprise search, biology research, and plugin management itself.

For this to actually hold up at enterprise scale, the book lists four enterprise-grade requirements: governance controls (an organization-specific plugin marketplace that turns governance from reactive shutdowns into proactive curation and distribution, stamping out shadow AI sprawl before it spreads); security by design (tasks run locally, nothing gets uploaded for cloud processing, resolving the most common objection to enterprise AI before it's even raised); auditability (compatible with OpenTelemetry, so you can see exactly when AI did what, on whose behalf); and integration and continuity (works inside your existing CRM and document systems, no context lost switching between Cowork and Excel, PowerPoint).
Four principles + a three-phase timeline, quoted in full
After all these stories about other companies, it comes down to you. The book's practical playbook comes in two pieces: four principles, each one pushing back against a common impulse; and a six-month timeline that puts the principles on the calendar. These two pieces are the most actionable part of the whole book, translated in full below — copy them directly and follow them.
① Start with specificity, not scale
Give Claude your standards, your tools, your institutional context from the beginning. Employees who receive generic output from a first interaction rarely give the tool a second chance. The organizations in this guide succeeded because they gave Claude enough context to produce output that felt like it came from someone who genuinely understood the business.
② Choose pilots with a measurable finish line
Each of the three pillars has its own success metrics: people gains are measured by adoption rate and time saved; process gains are measured by cycle-time compression and quality scores; product transformation is measured by revenue impact and speed to market. Define your success criteria before the pilot starts, so the results aren't ambiguous.
③ Build plugins for reuse from day one
The temptation is always to build a quick fix for one team and worry about reuse later. Resist it: a plugin built for one team should benefit the entire organization. Encode tribal knowledge once, and every team that installs that plugin gets the benefit immediately. The marginal cost of sharing a plugin is zero, while the marginal value is enormous.
④ Never underestimate the governance layer
Admin controls, auditability, and organization-specific marketplaces are prerequisites for broad rollout, not features you bolt on after adoption takes off. Organizations that skip governance early end up spending more time cleaning up unsanctioned usage than they saved by moving fast.
English original
• Start with specificity, not scale. Give Claude your standards, your tools, your institutional context from the beginning. Employees who receive generic output from a first interaction rarely give the tool a second chance. The organizations in this guide succeeded because they gave Claude enough context to produce output that felt like it came from someone who understood the business. • Choose pilots with a measurable finish line. Each of the three pillars has different success metrics. Smarter employees might be measured by adoption rates and time savings. Faster processes might be measured by cycle time compression and quality scores. Transformative products might be measured by revenue impact and speed to market. Define your success criteria upfront, before the pilot starts, so the results are unambiguous. • Build plugins for reuse from the beginning. The temptation is to build a quick solution for one team and worry about reuse later. Resist that temptation. Plugins built for one team should benefit the entire organization. When you encode tribal knowledge once, every team that installs the plugin gets the benefit immediately. The marginal cost of sharing a plugin is zero, while the marginal value is enormous. • Never underestimate the governance layer. Admin controls, auditability, and organization-specific marketplaces are prerequisites for broad rollout, not features you add after adoption takes off. The organizations that skip governance early spend more time cleaning up unsanctioned usage than they saved by moving fast.
Phase 1 · First few weeks: setting evaluation and success criteria
The first several weeks focus on exactly one thing: evaluation and success criteria. Identify two to three teams with clear pain points and measurable workflows; install relevant plugins from the open-source repository, or build custom plugins that encode your team's specific standards and processes; define what success looks like before anyone starts using the tool.
For example: for a sales team, that might be "call prep time cut by 50%"; for a legal team, "contract review turnaround compressed from five days to one"; for a documentation team, "first-draft quality reaching 80% of the final approved version." How specific your success criteria are matters a lot: a vague goal like "improve productivity" only produces vague results that are easy to dismiss.
Phase 2 · Months 2-3: launch a champion pilot
The second and third months are the champion pilot: two to three teams use Claude Cowork with configured plugins in real production workflows (not sandboxed experiments). Measure adoption weekly; collect qualitative feedback alongside the quantitative metrics, because the moments when employees discover unexpected value are often more informative than the time-saved calculations.
The goal of this phase isn't perfection — it's proof of value, plus a clear picture of what needs to change before broader rollout.
Phase 3 · Months 4-6: scaling impact
With a successful proof of concept in hand, months four through six shift to scaling and governance: deploy admin marketplace controls, establish plugin review and approval workflows, and roll out to more teams using the plugins and configurations refined during the pilot.
Every team that comes online benefits from institutional knowledge already encoded, so the second wave of adoption moves faster than the first, and the third wave moves faster still. This is the compounding dynamic actually running: every investment in context, configuration, and governance makes the next deployment cheaper and more effective.
English original
Phase 1: Setting evaluation and success criteria For the first several weeks, focus exclusively on your evaluation and success criteria. Identify two to three teams with clear pain points and measurable workflows. Install the relevant plugins from the open-source repository or build custom plugins that encode your team's specific standards and processes. Define what success looks like before anyone starts using the tool. For a sales team, for example, that might be call prep time reduced by 50 percent. For a legal team, it might be a contract review turnaround cut from five days to one. For a documentation team, it might be first-draft quality reaching 80 percent of the final approved version. The specificity of the success criteria matters: vague goals like "improve productivity" produce vague results that are easy to dismiss. Phase 2: Launching a champion pilot The second and third months of the initiative are the champion pilot. Two to three teams use Claude Cowork with their configured plugins in production workflows, not sandboxed experiments. Measure adoption weekly. Collect qualitative feedback alongside quantitative metrics, because the moments when employees discover unexpected value are often more informative than time-saved calculations. The goal of this pilot phase is not perfection but proof of value and a clear understanding of what needs to change before broader rollout. Getting started with Claude Cowork in the help center provides practical guidance for configuring access and managing this initial deployment. Phase 3: Scaling impact Successful proof of concept in hand, months four through six shift to scaling and governance. It's time to deploy the admin marketplace controls, establish plugin review and approval workflows, and begin the rollout to additional teams using the plugins and configurations refined during the pilot. Each team that comes online benefits from the institutional knowledge already encoded, which means the second wave of adoption moves faster than the first. The third wave moves faster still. This is the compounding dynamic in action: every investment in context, configuration, and governance makes the next deployment cheaper and more effective.
That line, "the second wave moves faster than the first, the third wave faster still," is exactly what the compounding mechanism from earlier looks like once it's on the calendar: it's not because the model got better — it's because the context, configuration, and expert feedback accumulated earlier are all doing the accelerating for you.
You don't need a perfect plan, you need a concrete starting point
We started from a salesperson's few seconds before the phone rang. By now you can probably see that what was hiding in those few seconds wasn't just a few saved hours.
Let's collapse the whole chain of logic: the model isn't the deciding variable — the evidence is that the same model produces wildly different results at different companies; the deciding variable is scope — whether you treat AI as a point tool or as the foundation for transformation. Point solutions are inevitably mediocre because they run on generic output, and generic output always "still needs work." Transformation runs on encoding organizational context into the system, and the numbers from the three pillars (L'Oréal's 99.9%, Lyft's 87% drop in resolution time, Rakuten's 97% drop in critical errors) repeatedly prove the same thing: what widens the gap is embedding depth and thickness of context, not the model itself. And this advantage compounds: expert feedback flows back into the knowledge base, agent memory accumulates, plugins turn institutional knowledge into shareable infrastructure — so the distance between first-movers and everyone else keeps growing.
The book calls out the most common, and most fatal, mistake: waiting until the strategy is complete before taking the first step. Successful organizations do the opposite — they cut in through a narrow opening, learn fast, then expand with conviction. You don't need a perfect plan. You need exactly three things: a concrete starting point, a set of quantifiable success criteria, and the willingness to keep learning from what actually happens next.

The AI adoption divide: treat it as a "tool" and you get mediocrity, treat it as "infrastructure" and you get compounding returns
Anthropic's enterprise rollout guide: four principles + a six-month, three-phase timeline, backed by real numbers from L'Oréal, Lyft, and Rakuten. One line: the gap isn't about how strong the model is — it's about how much of "your own company's context" you feed it.
↓ Read this one page · has animated diagrams
Enterprises are long past the "should we use AI" question — the question now is how. Most companies end up with "point solutions": a chatbot here, a summarizer there, impressive in a demo, but never scaling and never changing how the company actually runs. In short: they lined up a row of question-answering machines instead of letting AI break down its own steps and get the work done end to end.
✘ But it doesn't know your company's standards, your jargon, or the rules that only live in veteran employees' heads
So the first reaction to it is always "this still needs fixing" — fix it until it's usable. The gap isn't the model, it's that you never fed it your own company's context (standards, terminology, internal rules). What this guide is trying to solve is exactly this: how to get AI output that ships without edits.
The underlying mechanism, stated plainly: what makes AI actually useful is feeding it your company's own context, then looping expert-vetted feedback back into the system — so instead of starting from zero every time, it gets more accurate the more it's used. That's compounding (like compound interest — the more you use it, the more valuable it gets).
In practice, that's four principles, each pushing back against a common impulse:
- Start specific, not big: feed it your standards, tools, and jargon from day one — don't expect generic output to be usable as-is.
- Set a measurable finish line first: like "cut call prep time in half," not a vague goal like "improve productivity."
- Build plugins for reuse from day one: a plugin (a team's tools, processes, and materials packaged into one install-and-go bundle) shouldn't be a one-off patch job for a single team — the marginal cost of sharing it is close to zero.
- Set up governance guardrails early: who can use what, and whether it's logged — don't wait for something to go wrong and patch it after, the cleanup costs more than the time you saved by rushing.
With the principles in place, the guide also gives a six-month timeline that puts them on the calendar.
Both the four principles and the three phases are fully quoted with English originals in the article — copy them directly and follow them.
"How much faster" means nothing in the abstract — translate it into the same specific task taking wildly different amounts of time and it lands. (There's also a hard gate: in industries like finance and healthcare, customer data can't leave the company's "trust boundary" — the scope of data a customer is willing to hand over — or even the strongest product never ships.)
client brief
waits for AI to deliver.
still needs fixing
doesn't use the jargon.
all the company's know-how.
internal know-how
to "ships as-is."
keep getting faster?
the system remembers
doesn't start from scratch.
- × All self-reported by the companies
- × No third-party reproduction
- × The tool is their own Cowork
it's how much
of your own company's context you feed it.
you pull ahead by is a year of compounding
