Operations team reviewing dashboards on a knowledge AI platform linking connected department workspaces

Stripe’s teams built over 4,000 micro-agents with a no-code builder before anyone admitted the obvious: nobody could monitor or maintain them. Sales reps and finance analysts were writing near-identical prompts of wildly different quality, while others quietly repurposed coding agents and handed security a problem it never asked for. Then Stripe shipped Kai, its knowledge AI platform. Within two weeks of the April launch, most of the company was using it. Today 83% are weekly active users, including nearly all of go-to-market.

The interesting part is why the first two attempts failed. It wasn’t model capability. It was architecture, specifically how you scale expertise that lives in dozens of heads across your organization. If your quality and operations teams are drowning in manual knowledge work, that distinction decides whether your AI spend pays back.

Your Engineers Got AI. Your Quality Managers Got a Chatbot.

Stripe was blunt about the divide. Coding agents like Claude Code and Codex transformed engineering, while sales reps, finance analysts, and technical account managers “felt left behind by the AI wave.” Same story on the plant floor. Process engineers are running experiments with LLMs, and your quality manager is still copying deviation data between a spreadsheet and an email thread.

The reason isn’t budget or appetite. Coding work has a uniform shape: edit files, run tests, commit. A single agent architecture fits it. Knowledge work does not. Preparing a compliance review, triaging a customer complaint, and closing out a CAPA each need different data, different tools, and a different definition of done.

That mismatch is why generic assistants stall at novelty. People try them, hit the edge of what they know, and go back to the spreadsheet.

Split screen showing engineers using coding agents while quality managers lack a knowledge AI platform

What Stripe Actually Built With Kai

Kai handles the work that never fit a chatbot window. Querying data warehouses. Researching accounts before sales calls. Triaging incidents. Modeling revenue scenarios. Preparing compliance reviews. Five jobs with five different tool sets, five different data sources, and five different definitions of what finished looks like.

That range is the point. Stripe didn’t ship one assistant and hope people found uses for it. It shipped a platform that gives each function an agent shaped around how that function actually works.

The 83% adoption number and why multi-turn sessions matter more

Eighty-three percent weekly active is a strong number, and it will be the one your CFO quotes. It is not the most useful one.

The signal is in the session shape. Most Kai sessions require many turns, with users doing deep research, creating specific artifacts, or refining assets before sharing them internally or externally. People are producing deliverables, not sampling a novelty. One-shot Q&A usage tells you nothing about whether a tool has entered the workflow. Multi-turn artifact creation tells you it has.

Why Stripe treated this as infrastructure, not a tool purchase

Stripe named three requirements before writing code: scale expertise without centralizing it, meet users wherever they work, and enforce guardrails that don’t exist in code. None of those are features you buy. They are architectural decisions you commit to.

The people who know how to triage a billing escalation sit nowhere near the team building agent infrastructure. Stripe accepted that and built a system that models distributed judgment instead of trying to concentrate it. Buy a tool and you get one workflow. Build infrastructure and you get every workflow your experts can describe.

The Two Failure Modes Stripe Hit First (And Most Companies Are Living In Right Now)

Agent sprawl: 4,000 agents and no way to maintain them

Stripe’s no-code builder worked exactly as intended. Anyone could ship a workflow-specific agent with tool access, and thousands of people did. The problem showed up later: a proliferation of micro-agents that became, in Stripe’s words, “increasingly hard to monitor and maintain.”

Manufacturing organisations hit the same wall at smaller scale. One quality engineer builds a custom GPT for CAPA drafting. Another builds a near-identical one for supplier corrective actions. A third writes better prompts than both, but nobody knows that, so nobody copies them. Six months later you have forty half-working tools, no owner, no version control, and no idea which ones are feeding auditors wrong answers.

Shadow AI: when non-engineers borrow engineering tools

The second failure mode is quieter and more expensive. At Stripe, some non-engineers gave up waiting and moved their workflows into coding agents. Those tools are genuinely powerful. They also triggered security concerns and created a support burden for code quality teams that had never supported non-engineers before.

On the plant side this looks like a validation engineer pasting batch records into a consumer chatbot because it is faster than the approved route. Or an ops manager running production data through a tool procurement never reviewed. The intent is fine. The exposure is not, particularly under ISO 13485, IATF 16949, or any regulated process where data lineage matters.

Both failure modes share a root cause. People with real expertise were handed tools built for a different job, then left to improvise governance themselves.

Side-by-side diagram contrasting duplicate agent sprawl with one overloaded bot before a knowledge AI platform

The Three Design Principles That Made Kai Stick

Stripe boiled its requirements down to three: scale expertise without centralizing it, meet users wherever they already work, and enforce guardrails that don’t exist in code. The middle one is the easiest to solve and the one most vendors sell you. The other two decide whether the thing survives past month three.

Distributed expertise: the domain owner writes the workflow, not the AI team

Stripe’s own description of the problem is exact: the people who know how to triage a billing escalation “are not centralized and certainly aren’t on the team building the agent infrastructure.” Swap billing escalation for a supplier nonconformance or an out-of-spec batch investigation and nothing changes. Your validation lead knows what a defensible deviation looks like. Your IT team does not.

So stop routing every workflow through a central build queue. The domain owner should be able to define the steps, the data sources, and what “done” means, inside a structure someone else maintains. If encoding a new workflow requires a ticket and a sprint, you have built a bottleneck with an AI logo on it.

Guardrails for knowledge work, where there’s no test suite to catch mistakes

Code has an advantage nobody talks about. Tests fail loudly. A wrong CAPA rationale or a misread spec limit passes quietly into a document that an auditor reads eighteen months later.

That means the controls have to be structural, not aspirational. Restrict which systems each agent can read from. Require citations back to the source record. Log every session so a reviewer can reconstruct how an output was produced. In regulated manufacturing, traceability is not an add-on feature, it is the precondition for using the output at all.

What This Playbook Looks Like on a Factory Floor

The manufacturing equivalents are obvious once you look for them: CAPA investigations that need every historical nonconformance on a similar failure mode, supplier audit prep, deviation triage, batch record review, ISO or IATF evidence gathering, shift handover summaries that currently live in a notebook. Each needs different data, different tools, and a different definition of done.

Pick the workflow before you pick the platform

Start with one workflow that happens weekly or daily, causes visible friction, and has a single named owner. CAPA research is usually the best first candidate. The quality engineer already knows what a good investigation looks like, the source data sits in your QMS, and the output is a document someone reviews.

Do not build platform architecture for one workflow. Stripe only needed a knowledge AI platform after thousands of agents made maintenance impossible. Build your second and third workflows, then look at what plumbing they share: the same document repository, the same audit trail, the same approval step. That shared layer is what you standardise.

Calculating hours returned and what to do with them

Measure the current state before you change anything. Time three recent CAPA investigations end to end, separating research from analysis and writing. Same for audit prep and batch record review. You want hours per week per role, not a vague percentage, because that number is what your finance director will argue with.

Then decide where those hours go before they get absorbed. Returned time defaults to more of the same reactive work unless someone redirects it. Point it at the things quality teams never reach: trend analysis across suppliers, process capability studies, preventive action that actually prevents. That redeployment is the return, not the headcount you did not hire.

Quality engineer on factory floor reviewing CAPA records in a knowledge AI platform

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

The 2026 Split: Companies With an AI Platform vs. Companies With AI Tools

Stripe’s numbers did not come from a better model. Everyone has access to the same frontier models. What Stripe did differently was treat internal AI as infrastructure: one platform, ownership pushed out to the people with the domain knowledge, access embedded in the tools those people already open, and guardrails enforced centrally instead of hoped for.

That structure compounds. Every workflow someone encodes makes the next one cheaper to build, because the tool connections, permissions, and audit trail already exist. Buying a point tool per department does the opposite. Each purchase adds a login, a data agreement, a separate prompt library, and one more thing nobody maintains after the champion changes roles.

Note also what Stripe reported about usage: most Kai sessions require many turns, with users doing deep research or refining artifacts before sharing them. That is not chatbot behaviour. If your AI usage is single-question-single-answer, you are not doing knowledge work with it yet.

Your 90-day starting move

Start by finding out what is already happening. Somebody in quality is pasting nonconformance text into a consumer chatbot right now. Ask, without consequences attached, and you will get an honest map of where the manual load actually sits. Shadow usage is free requirements-gathering.

Then rank your knowledge workflows by hours consumed per month, not by how interesting they look. Take the top three. Those are your candidates, and the hour count is your ROI baseline before anyone signs anything.

Last, name the person who owns guardrails: what data agents can touch, what they can send outside the company, who reviews outputs before they reach a customer or an auditor. Decide that in the first 90 days or it gets decided for you, badly.

Source: stripe.dev

Leave a Reply