Mathematician at a chalkboard buried under stacks of AI generated contributions awaiting human review

Thomas Bloom built erdosproblems.com in May 2023 to promote a few hundred hard math problems to human readers. By late 2026 the site held 1,221 problems, more than 9,000 comments, and up to 25,000 daily visitors. Then came the wave he describes as AI-produced solutions provided with no explanation: technically correct, useful, and quietly pushing out the people who actually wanted to think about the mathematics. The governance never caught up with the volume.

Your operation is heading for the same squeeze. When AI generated contributions start arriving faster than your team can review them, the bottleneck moves from producing work to trusting it. Below, we break down what Bloom’s response reveals about handling that shift, and the review structures you need in place before the flood, not after.

A Maths Website With 9,000 Comments Just Hit a Wall AI Built

On 6 October 2026, Thomas Bloom published a post announcing changes to his site. His framing of the speed is worth sitting with: in mathematics, AI has gone from “essentially useless to helping solve some of the hardest problems in mathematics in less than a year.” He never set out to build a benchmark. He built a reference for humans, and a large pool of simply stated, unsolved questions turned out to be the ideal showcase for machines.

The results cut both ways. Real answers arrived. So did a second effect Bloom names directly: mathematicians who stopped thinking about these problems because they believed they could not compete.

Nothing about that dynamic is specific to number theory. It shows up wherever expert judgment is the product.

Screenshot of Thomas Bloom's erdosproblems.com post announcing new rules for AI generated contributions
Photo by Vitaly Gariev on Pexels

What Actually Broke: Volume, Opacity, and Discouraged Experts

Three distinct things went wrong on the site, and Bloom names each of them. They are not the same problem wearing different hats. Each one maps cleanly onto something your operation will face, and each needs a different fix.

The second failure is reputational. Once machines could touch Erdős-style problems, some mathematicians began writing the whole area off as “easy/recreational.” The work did not get easier. The perception shifted because a machine could do part of it, and that perception drags down the people still doing the hard parts. Watch for the same drift when your deviation investigations or CAPA drafts start getting AI assistance.

The third is withdrawal. Bloom reports that many have simply stopped thinking about these problems, believing they cannot compete. That is the most expensive outcome on the list, because it removes capability you cannot buy back quickly.

Answers without explanation create verification debt

An unexplained correct answer is still work for someone. Somebody has to check it, and checking a result you did not derive is often slower than deriving it yourself. Multiply that across a queue and you have a backlog of unverified output sitting in your system, labelled as progress.

In a plant, this looks like a root cause analysis that reaches the right conclusion with no visible reasoning chain. Your auditor will ask how you got there. “The model said so” is not an answer that survives a regulatory review, and it teaches your engineers nothing they can reuse next quarter.

The silent majority nobody measures

Bloom made a point worth stealing. The feedback he valued most came from people who had learned a lot from the site and never once commented. He calls them the large silent audience, and he built the changes around not drowning them out.

Your equivalent exists. They are the line supervisors and quality engineers who read every report, never push back in meetings, and quietly stop contributing when the queue fills with output they did not write and cannot interrogate. Nothing in your dashboard will flag it.

Bloom’s Governing Principle and Why It Beats a Blanket AI Ban

Bloom did not reach for a rule first. He reached for a principle. He borrowed it from Po-Shen Loh’s essay Why Do We Need Human Mathematicians Anymore?, which proposes a simple axiom:

We (humans) should help humanity flourish.

Bloom narrowed it to his own situation: the site should help the Erdős-community of humans, defined as those who are interested in Erdős-style mathematics and want to think about and understand it, flourish. That sentence is doing real work. It names a specific group, states what they need, and makes every policy question downstream of that answer.

Notice what he did not do. He did not ban machine output, which would have thrown away genuine results. He did not wave everything through either. He wrote down who the system serves, then let that decide the rules.

Define the humans your process exists to serve, then set the rules

Most companies skip this step and land on one of two defaults. The first is prohibition: a policy memo saying no AI in deviation reports or supplier assessments. Nobody enforces it, people use the tools anyway, and now the usage is invisible. You have lost oversight and gained nothing.

The second default is open season. Any tool, any team, no review. Quality improves on paper for about two quarters, then you notice the engineers who used to explain why a root cause analysis landed where it did have stopped explaining anything. Institutional knowledge drains out quietly, and you only find the hole when something goes wrong and nobody can reconstruct the reasoning.

Write your version of Bloom’s sentence before you write a single rule. Who is your process actually for? Probably the quality engineers who need to understand failure modes well enough to prevent the next one, not just close the current one. Once that is on paper, your AI contribution policy stops being a philosophical debate and becomes an operational test: does this input help those people get better at the work, or does it replace the thinking that makes them good?

A moderator reviewing flagged AI generated contributions on screen beside a crossed-out ban sign

The Same Flood Is Coming for Quality Reports and CAPA Logs

A model can draft a root-cause narrative, a deviation write-up, a supplier risk assessment, and three process improvement proposals before your morning stand-up ends. All of them will read well. None of them will carry a reasoning trail unless you demanded one, and a quality manager reviewing forty of these a week will stop reading carefully by Wednesday.

The failure mode is the one Bloom watched play out: output that is technically useful but opaque, and experienced people quietly stepping back because the machine filed first. Your best process engineer does not argue with a finished CAPA. They just stop writing them.

Label the source, keep the reasoning

Tag AI involvement at the point of entry, not during review. A field in the form, filled in before submit, with three states: human-authored, AI-assisted, AI-generated. Retrofitting that label later never happens, and audit trails built after the fact are worth nothing.

Then require the reasoning alongside the conclusion. Which batch records were read, which assumptions were made, what was ruled out and why. A root cause with no visible path to it is a guess that happens to be formatted correctly. Reject it the same way you would reject a deviation report with a blank investigation section.

What this protects during an audit

An auditor does not ask whether AI wrote your CAPA. They ask who concluded this, on what evidence, and why the corrective action follows. If the answer is a workflow step that auto-approved at 2am, you have a finding. Sign-off must attach to a named person who can defend the logic out loud, under questioning, months later.

That single rule does most of the work. It caps how many AI generated contributions can enter your system, because a named human can only defend what they actually understood. Volume gets throttled by accountability instead of by a review queue nobody has time for. Your engineers stay in the loop because their name is on it.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

Designing for Flourishing Experts, Not Just Faster Output

Bloom’s redesign is a bet, and it is worth naming as one. He is choosing a smaller, engaged community over a larger pile of answers. He noticed something most operators never measure: a large group of people who learned from the site, found new problems to work on, and never left a single comment. Those people do not show up in any throughput metric, and they are the first to disappear.

The same logic holds inside a plant or a quality function. The engineer who reads every deviation report and develops a feel for which supplier excuses are real is building capability that no dashboard records. Automate the drafting and that person gets more time to build it. Automate the judgement and the capability stops forming, which you will discover two years later when nobody can tell a plausible root cause from a correct one.

A three-question audit for any AI-assisted workflow

Run every proposed automation through three questions before it goes live. They take ten minutes and they catch most of the damage early.

  • Is this task low-value, or just hard to staff? Low-value work is safe to automate. Work you are automating because you cannot hire for it is judgement work wearing a disguise, and handing it to a model buys relief now and a skills gap later.
  • Can a reviewer see the reasoning, not just the conclusion? If the output arrives as a verdict with no trail, your reviewers will rubber-stamp it by the fortieth one.
  • Who gets better at their job because of this? If the honest answer is nobody, you have bought speed and sold expertise.

Make the second list, the hard-to-staff one, and treat it as a hiring and training problem that AI can support but not replace. That list is where your competence lives. Protect it deliberately, or the flood will decide for you.

Source: erdosproblems.com

Leave a Reply