Sam Altman said society will have to “accept some bad things” for AI to deliver its benefits. That framing made a lot of people uncomfortable. It shouldn’t. You already run that calculation every shift: an AQL sampling plan accepts a known defect rate, a maintenance interval accepts a known failure probability, a release gate accepts a known escape risk. Nobody calls that reckless. They call it engineering.
The problem is that almost nobody applies the same discipline to AI. Models get deployed into inspection, documentation, and scheduling workflows with no defined tolerance for what a wrong output costs. Below, we break down how to set an explicit AI risk tolerance per use case, where to put human checkpoints, and how to price the errors you’ve decided to live with.
Your Plant Runs on Defined Failure Rates. Your AI Pilot Doesn’t.
Walk the floor and every tolerance is written down. Scrap allowances sit in the budget. Planned downtime sits in the schedule. Someone signed their name to each of those numbers, and if actual performance drifts outside the band, there is a process that triggers. That is normal industrial practice, not a compromise.
Then an AI vision tool or a documentation assistant comes up for approval and the standard flips. Either the tool must be flawless before anyone will touch it, or it goes live with no stated error budget at all and nobody notices until a customer does. Both failures come from the same gap.
The argument here is narrow. Deciding whether to accept bad outcomes is settled. Deciding which ones, how many per thousand, and who answers for them is what almost nobody has documented.

What Altman Actually Said, and Why the Framing Is Self-Serving
Speaking to the Guardian on 5 October, Altman argued that society will need to tolerate a degree of harm in order to get the benefits AI can deliver, pointing to technologies we already live with and accept risk from. Cars kill people. Electricity kills people. We regulate, we engineer, we keep driving and keep the lights on.
That comparison is doing two jobs at once. One of them is honest. The other quietly shifts the bill to someone who never agreed to pay it.
The argument that holds up: utility and risk are never decoupled
No technology with genuine industrial utility has ever been risk-free. Compressed air, overhead cranes, robotic cells, chemical cleaning agents: every one of them carries a hazard profile, and every one of them earned its place because the upside was worth managing the hazard. Demanding zero risk from AI is not caution. It is a decision to forfeit the upside while competitors take it.
Quality leaders understand this instinctively. You do not refuse to run a process because it has a Cpk below infinity. You characterise the variation, set limits, and control it. The same logic applies to a model that classifies defects or drafts deviation reports. The question is never whether it will be wrong, it is how often, how badly, and what catches it.
The part that doesn’t: who gets to define ‘acceptable’, and who pays
“Society will have to accept some bad things” has no subject. Nobody in that sentence is named as the decision-maker, and nobody is named as the one absorbing the cost. In practice the gains concentrate with a handful of model vendors while the consequences land on customers, operators, and regulators who were never consulted.
That matters on your site. An operator accepting a known failure rate on their own balance sheet is engineering. A vendor accepting a failure rate on your balance sheet is a transfer of liability dressed up as progress. Your AI risk tolerance has to be defined by the party holding the recall, the audit finding, and the customer complaint.
Where the Tradeoff Logic Works in Operations, and Where It Collapses
The line is not about how important the task is. It is about whether a wrong answer announces itself. Ask one question before any deployment: can you detect the error before it becomes expensive? If yes, you can run an error budget. If no, you are gambling with someone else’s money.
Reversible, reviewed, low-stakes: where an error budget makes sense
Drafting a CAPA report, summarising audit findings, triaging supplier correspondence, pre-screening deviation records. In all four, a human reads the output before it does anything. The error surfaces during review, costs a few minutes of correction, and leaves no trace downstream.
Work the arithmetic. If a model drafts a deviation summary and gets one line in twenty wrong, the quality engineer fixes that line and still saves twenty minutes per record. That is a good trade, and it holds as long as the reviewer is genuinely reviewing rather than clicking approve. Measure the catch rate, not just the model accuracy. Reviewer fatigue is the real failure mode here, and it is easier to monitor than people expect.
Silent failures and regulated records: where zero tolerance is the correct answer
The logic collapses the moment AI output becomes the record. Batch disposition, release decisions, anything feeding a regulatory submission, anything that writes directly to a validated system. A wrong answer in those places does not get caught by a reviewer. It gets caught by an auditor, a recall, or a customer complaint eighteen months later.
Compounding failures deserve the same treatment. A model that mis-assigns root cause codes creates a trend analysis that is quietly wrong, and every decision built on that trend inherits the error. Same with planning tools that feed material requirements without a check loop. Set AI risk tolerance to zero in these workflows, or restructure them so the model advises and a named person decides. Advisory is fine. Autonomous is not, and no accuracy figure changes that.

Write Your AI Error Budget Before You Write the Business Case
An error budget is a one-page document. It takes an afternoon. Most teams skip it because it feels like paperwork, then spend six weeks in approval meetings arguing about whether the tool is “ready” with nobody able to say what ready means.
The five questions that turn ‘is it safe?’ into a signed-off threshold
- What are the specific failure modes?: Not “it might be wrong.” Name them. Misclassified severity, wrong product family, hallucinated root cause reference.
- What does each one cost, and how often?: Rough order of magnitude is fine. A misrouted minor NC costs an hour. A missed critical costs a customer notification.
- Who checks, and at what point?: The checkpoint has to sit before the irreversible step, not after.
- What number triggers rollback?: One threshold, measurable weekly.
- Whose name is on it?: One person, not a committee.
Filled in for AI-assisted non-conformance classification: target 92% agreement with the quality engineer on category and severity. Every record routed as minor gets auto-approved; every record the model flags as major or critical goes to a human regardless of confidence. Ongoing verification: 30 records sampled weekly against engineer re-review. Below 88% for two consecutive weeks, classification reverts to manual and the quality manager owns the recovery plan.
Monitoring drift: why your day-one accuracy number expires
Validation accuracy describes the data you had, not the data you will get. New product introductions, a changed supplier, a reworded defect taxonomy, a different shift writing the descriptions. Each one moves the input distribution and your model does not tell you it happened.
Track the sampled agreement rate as a control chart, not a one-off report. You want the trend, because drift usually shows up as three mediocre weeks before it shows up as a failure. Review it in the same meeting where you review scrap and first-pass yield.
Teams that define this upfront ship faster. Approval stops being a debate about comfort levels and becomes a decision about whether 92% beats the current manual baseline.
Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.
The Companies That Win Will Be the Ones That Quantified the Downside First
Through 2026, model capability will keep moving faster than any governance function can document it. That gap is permanent now. Waiting for the tooling to stabilise before you write policy means you will never write policy.
The competitive split will not be adopters versus holdouts. Almost everyone is adopting something. The split will be between organisations that can state a number when an auditor asks what error rate they accept, and organisations that answer with a paragraph about human oversight and continuous improvement. One of those answers survives a certification audit. The other triggers a finding.
Watch what the quantified group actually does differently. They move from proposal to production in weeks, because the argument about readiness was settled on paper before anyone asked for budget. They also kill bad ideas faster, which matters more. When the projected cost of failure exceeds the projected savings, the decision takes ten minutes instead of two quarters of pilot extensions.
The unquantified group splits into two failure modes, both expensive. Some stall indefinitely, running a fourth evaluation on a tool whose value was obvious a year ago. Others ship fast and then discover, during a customer audit or a recall investigation, that nobody can produce evidence of what the system was permitted to get wrong. The second outcome is worse. Pilot purgatory costs you time. An undefensible deployment costs you a qualification.
Here is the practical move. Pick one use case this quarter, not three. Choose something reversible with a human already in the loop, because you want the first attempt to be easy. Write its error budget on a single page: failure modes, cost per occurrence, acceptable frequency, detection method, and the person who signs it.
Then treat that page as your template. The second one takes an hour. The tenth takes twenty minutes, and by then your stated AI risk tolerance is a documented standard rather than a standing agenda item nobody closes.
Source: theguardian.com