Your vision system flags a weld as defective. It is 94% confident. It is also wrong, and it will never tell you that, because it answers every question the same way: fast, cheap, and without a second thought. That single blind spot kills more AI pilots on the factory floor than bad data ever will. In October 2021, a team from IBM Research and collaborators including Francesca Rossi, Murray Campbell and Lior Horesh published “Thinking Fast and Slow in AI: the Role of Metacognition”, borrowing Daniel Kahneman’s System 1 and System 2 framing to give machines a capability they were missing: knowing when to stop guessing and start reasoning.
Below, we break down what AI metacognition actually means in production terms, where it changes quality and maintenance decisions, and how to tell whether your current systems have it.

Your AI Gives the Same Confident Answer to Easy and Hard Questions
A textbook surface scratch and a hairline crack hidden under coating arrive at your classifier in the same format, and they leave it in the same format: one label, one probability, one millisecond. The model has no way to signal “this one is outside anything I have seen.” Your operators learn that quickly. Once they stop trusting the confidence score, they stop using the output, and the pilot dies without anyone formally killing it.
Bergamaschi Ganapini and her co-authors name the underlying condition precisely. We are “still mostly seeing instances of narrow AI,” systems whose wins come from “huge datasets and computational power” rather than any understanding of their own limits.
Retraining does not fix this. The missing layer is metacognition, not accuracy.

What the SOFAI Paper Actually Proposes: Fast Agents, Slow Agents, One Arbiter
The ten authors (Bergamaschi Ganapini, Campbell, Fabiano, Horesh, Lenchner, Loreggia, Mattei, Rossi, Srivastava and Venable) do not propose a bigger model. They propose a multi-agent AI architecture where incoming problems get routed. Some go to System 1 agents. Some go to System 2 agents. Both draw on a shared model of the world, holding domain knowledge about the environment, and a model of “self”, holding information about past actions and the skills of each solver.
That distinction matters when you are buying. A vendor selling you a larger model is selling more of the same reflex. An architecture that can route is selling judgement about which reflex to use.
System 1 agents: pattern recall at production speed
System 1 agents “react by exploiting only past experience”. No search, no deliberation, no alternative hypotheses. They see a pattern they have seen before and return the answer that worked last time.
This is the right tool for the vast majority of your volume. Inline inspection, barcode reads, standard part classification. You do not want reasoning latency on a line running at cycle time. Keep System 1 fast and cheap, and keep it inside the boundary of what it has actually seen.
System 2 agents: deliberate search, activated on purpose
System 2 agents are “deliberately activated when there is the need to reason and search for optimal solutions beyond what is expected from the system 1 agent”. The activation is the whole idea. Something decides that fast recall is not good enough here.
The cost is real: search takes time and compute. So you spend it selectively, on the ambiguous defect, the novel failure mode, the case where the model of self says this solver has a weak track record.
The Two Models That Make the Handoff Work: World and Self
Routing is only as good as the information the router has. The paper gives both agent types two supports: a model of the world and a model of self. Most vendors will tell you they have the first one. Almost none have the second.
Why a model of self is the escalation trigger
The model of self holds “information about past actions of the system and solvers’ skills.” Read that operationally. The system keeps a running record of where it has been right, where it has been wrong, and which solver handled each case.
That record is what makes escalation possible. A classifier with no memory of its own performance cannot distinguish a part type it has judged correctly ten thousand times from one it has seen twice. A system with a model of self can, and it can hand the second case to a slower agent that actually reasons. Without this, AI decision confidence is a number the model produces about itself with no evidence behind it.
What domain knowledge has to contain to be usable
The model of the world is domain knowledge about the environment, and in manufacturing that means specifics, not a generic ontology. Tolerances per part family. Which defect classes are safety-critical and which are cosmetic. Process constraints that make certain classifications physically impossible.
When that knowledge is missing, System 2 has nothing to reason over. It searches for an optimal solution in a space it cannot evaluate. So the practical question during vendor evaluation is simple: where does this system store what it knows about my plant, and where does it store what it knows about itself? If the answer to the second is a confidence threshold in a config file, you do not have metacognition.

Metacognitive Routing vs. Always-On Reasoning: The Cost Argument
Three architectures are on the table in 2026. One fast model handling every input. One heavyweight reasoning model handling every input. Or a router that decides, per case, which one gets the job.
| Approach | Latency on the line | Compute per decision | Auditability |
|---|---|---|---|
| Fast model everywhere | Milliseconds, fits takt time | Negligible | One label, no reasoning trace |
| Reasoning model everywhere | Seconds, breaks takt time | High and constant | Full trace, mostly on trivial cases |
| Metacognitive router | Fast by default, slow on exception | Spent where uncertainty is real | Trace exactly where you need one |
Where always-fast quietly costs you scrap and rework
A single fast model never escalates, so its errors do not surface as errors. They surface as scrap two stations downstream, or as a customer return six weeks later. By then the link back to the classification decision is gone.
Worse, your CAPA process has nothing to investigate. There is no record of deliberation because no deliberation happened. You are auditing a probability score, which tells you what the model output and nothing about why.
Where always-slow costs you throughput and budget
Running deliberation on every input is the opposite mistake, and it is expensive in two currencies. Inference cost scales linearly with volume. Latency scales with it too, and a line running at a few seconds per part cannot wait on a reasoning chain for a part that is obviously fine.
The SOFAI authors are specific on this point: System 2 agents are “deliberately activated when there is the need to reason.” Need. Not always. That word is where the entire cost argument for AI metacognition sits.
Applying Fast/Slow Routing to Quality and Operations Workflows
Start with your decisions, not your models. Every quality and operations workflow contains two populations: high-volume routine calls (inspection pass/fail, document classification, spec and tolerance lookups) and low-volume consequential calls (borderline defects, deviation investigations, supplier non-conformance root cause). The first population wants speed. The second wants a System 2 agent that, in the paper’s words, is “deliberately activated when there is the need to reason and search for optimal solutions.”
A three-step audit to map fast vs. slow decisions
- Count and cost: pull 90 days of decisions from one line or one process. Record volume, cycle time, and the cost of being wrong. Anything cheap to redo stays fast.
- Name the System 2 endpoint: for each slow category, write down exactly who or what takes the case. A reasoning model, a constraint solver, or a named quality engineer. If the answer is “someone notices eventually,” you have no escalation path.
- Set the trigger: confidence below threshold, input outside known distribution, or conflicting sensor and document evidence.
Then instrument the model of self. Log every decision, the confidence attached to it, the route it took, and the outcome once it is known. Calibrate the threshold against that log quarterly. Guessing a number at go-live and never revisiting it is how thresholds drift into uselessness.
ROI: fewer escalations to engineers, fewer missed edge cases
Two lines move. Routine volume stops reaching engineers, because the fast agent handles it and the router proves it should. Hard cases stop slipping through as confident errors, because they are flagged before they reach a shipment.
Track both: escalation volume per engineer per week, and escaped defects traced back to an automated call. If only one improves, your threshold is set wrong in a direction you can now measure.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.
What to Ask Your AI Vendor Before the Next Pilot
Every reasoning product on the market now sells some version of fast and slow. The vocabulary spread. The architecture mostly didn’t. What you get in a demo is usually a single model with a longer thinking budget, not a system that knows which mode a given case deserves.
Four questions will separate the two in under an hour:
- How does the system decide it doesn’t know? Ask for the mechanism, not the number. A probability threshold is not self-knowledge.
- What does it log about its own past performance? Per solver, per case type, over time. If the answer is “we log predictions,” there is no memory of skill.
- What happens on escalation? Who or what receives the case, what extra context travels with it, and how long the slow path takes.
- Does the escalation change future routing? If corrected cases don’t alter which agent gets the next similar input, nothing is learning.
The self-knowledge question most vendors can’t answer
The second question is where demos stall. Most vendors have built the model of the world, the domain knowledge about the environment, because that is what training data produces. The model of self, “information about past actions of the system and solvers’ skills,” requires deliberate instrumentation nobody asked for. Five years after the 2021 submission, that gap is still the norm.
You can specify it yourself. Require per-decision logging of solver identity, confidence, escalation status, and eventual ground truth, written to a store you own. That log is the judgment layer, and it is the part competitors cannot buy off a shelf. The model is a commodity. What sits around it is yours.
Source: arxiv.org