Erich Grunewald argues you should almost never use AI to write anything substantive. Not a blog post, not a research report, not a memo. His reasoning: the writing process is the thinking process, AI prose is vague and wrong in ways that are hard to catch, and passing it off unlabeled is misleading. He is not anti-AI. He explicitly endorses it for transcription, data analysis, search, brainstorming, and feedback on drafts. What he objects to is delegating the sentence-level work where the thinking actually happens.
That distinction matters more on a factory floor than on Substack. A CAPA that reads well but reasons badly will pass internal review and fail an audit. This article maps which parts of your quality documentation AI should own, which parts your people must keep, and how to tell the difference before a regulator does.
Your CAPA Report Was Written by a Model That Didn’t Investigate Anything
Walk into most quality departments and you will find the same workflow. An engineer dumps six bullets into a chat window, gets back a fluent deviation report, skims it, ships it. Same for root cause narratives. Same for audit responses that go to a notified body.
The output reads better than what the engineer would have written. That is exactly the problem. Grunewald’s point is that writing is where the thinking happens, and Paul Graham puts it more bluntly:
Writing about something, even something you know well, usually shows you that you didn’t know it as well as you thought. Putting ideas into words is a severe test.
A CAPA is that test. The sentence where you explain why the seal failed is the sentence where you discover you don’t actually know. Skip it and you have a document that passes review and a root cause that is still live on the line.

Grunewald’s Three Objections, Translated to the Factory Floor
Objection one: thinking happens in the sentences. Clara Collier describes it well when she says there is no part of the writing process in which she is not thinking and changing her mind. A root cause analysis works the same way. You start convinced it was operator error and three paragraphs in you realise the work instruction was revised in March and nobody retrained the second shift.
Objection two: the prose is vague and wrong in ways that are hard to notice. Objection three: unlabelled AI text is misleading to the reader. Both land harder on a supplier quality assessment or a management review summary than they do on a personal essay.
Why ‘wrong in hard-to-notice ways’ is a compliance risk, not a style problem
A model fills gaps with plausible connective tissue. It writes “the deviation was contained within the affected lot” because that sentence pattern appears in thousands of similar documents, not because anyone checked the containment records.
Graham calls putting ideas into words a severe test. Generated text skips the test entirely. The errors that survive are the ones that read correctly, which means your reviewer approves them and they sit in the quality system until an investigator pulls the file.
The credibility cost when your auditor can tell
Auditors read hundreds of these documents. They know what a hedged, over-fluent, evidence-light narrative looks like, and they have started asking follow-up questions specifically to test whether the author understands their own report.
Once an auditor suspects the CAPA narrative was generated rather than investigated, the scope of the audit widens. You are no longer defending one finding. You are defending the integrity of the whole system.
The Tasks Grunewald Explicitly Says AI Should Do
The list is specific: transcribing audio, analyzing data, searching for information, brainstorming, and giving feedback on drafts. Line and copy editing too, along with rewriting a passage to make it clearer or tighter. That is a large amount of work to hand over, and most quality teams are handing over almost none of it while happily handing over the part they should keep.
Map it to what your team actually does. AI transcribes the shop-floor interview you recorded during the deviation investigation, including the operator’s asides you would have missed. It pulls every non-conformance coded to the same work centre over eighteen months and tells you which shift and which supplier lot cluster together. It retrieves the three prior CAPAs that touched the same equipment. It reads your draft root cause and tells you where the causal chain skips a step.
Retrieval, extraction, transcription, trend analysis, first-pass critique. AI owns those. You own the argument.
The ‘deliberate accept or reject’ rule as a deployment standard
Grunewald attaches one condition to the editing use case: it is fine “as long as all the edits are deliberately accepted or rejected by a human.” That is a throwaway line in an essay and a deployment standard in a regulated plant. Bulk-accept is the failure mode. If your tooling offers an “apply all suggestions” button for anything that enters the quality system, turn it off.
Build it into the workflow rather than the training deck. Track changes on, one decision per edit, reviewer initials against the record. Auditors do not care that a model suggested the wording. They care whether a named person read it, understood it, and chose it.

Where the Line Actually Sits: A Task-by-Task Split for Quality and Ops
Here is the test, and it takes about four seconds to apply. If you would change your mind while writing the document, don’t delegate it. If the conclusion is already fixed and the writing is just transcription, automate the whole thing.
That single rule sorts almost every document your quality function produces.
High-judgement documents: AI as researcher and critic
Root cause analyses, risk assessments, change control justifications, supplier disqualification rationales, strategy memos to the board. In each of these, the argument gets built as you write it. You commit to a position in paragraph two and the evidence in paragraph five forces you to retract it. That’s the work.
AI still earns its place on these. Have it pull every prior deviation touching the same equipment, stress-test your fault tree, argue the opposite conclusion back at you. Grunewald’s own carve-out covers line and copy editing, “as long as all the edits are deliberately accepted or rejected by a human.” Accept or reject each one. No bulk approvals.
Low-judgement documents: AI as the whole pipeline
Standard work instructions generated from a validated template. Shift handover summaries built from structured MES logs. Meeting minutes from a transcript. Training record confirmations. Batch release cover sheets where the data already passed its checks.
The thinking here happened months ago, when someone designed the template or set the acceptance criteria. Writing is pure formatting. Hand it over completely, with a light verification step rather than a full review, because a human rereading generated minutes adds nothing except delay. Most teams have this backwards: heavy review on documents that were never going to teach them anything, and a quick skim on the ones that would have.
What This Costs You If You Get It Backwards
The two mistakes are not symmetrical. Automate the wrong low-judgement document and you lose a formatting pass. Automate the wrong high-judgement document and you lose the finding itself, and you don’t find out for eighteen months, when a notified body pulls the file.
Grunewald concedes the upside plainly: using AI to write is “less effortful and much faster than writing yourself.” That is true for a batch record summary and equally true for a root cause analysis. The difference is what you are buying with the saved time. In one case, hours. In the other, a corrective action that closes cleanly, changes nothing on the line, and lets the same deviation recur under a new number.
The slower cost is skill. A quality engineer who has never written their own investigation narrative has never had the experience of changing their mind halfway through one. After two years of that, you have a department that can approve documents and cannot reason about them. No audit will catch this. Your recurrence rate will.
Audit what you are already delegating. Pull ten documents your team produced last quarter and ask:
- Who wrote the first draft: a person with the evidence in front of them, or a model with six bullets?
- Did the conclusion change during drafting: if no document ever shifts its position mid-write, the thinking is happening somewhere else, or not at all.
- Would it survive an auditor asking “why this cause and not that one”: fluency is not traceability.
- Is the AI involvement recorded: unlabelled AI text in a regulated file is a data integrity problem before it is an ethics one.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.
Building an AI Policy That Survives the Next Model Release
Grunewald is careful to limit his claim. He says he is talking about “the AI models that exist now and that I expect to exist in the near future,” and concedes that better models may eventually earn the writing. He adds a sharp caveat: at that point it probably makes more sense to delegate the whole research process, because a model good enough to write the analysis is doing the thinking too.
That caveat is the reason your policy should never name a model. Write a rule around GPT-5 or Claude and it expires the next time a vendor ships. Write it around judgement density, meaning how much of the conclusion gets formed while the document is being drafted, and it holds regardless of what arrives next quarter. When a model genuinely crosses the line, you revisit the artefact list, not the entire policy.
Disclosure as a default, not an exception
Most quality teams treat disclosure as something you do when someone asks. Flip it. Every document carries a line stating what the model did: transcription, first-pass summary, editing, nothing. Auditors do not object to AI involvement. They object to discovering it after the fact, in a document that claimed a human investigator reached the conclusion.
Three things belong in writing, and they fit on one page. First, the disclosure rule, applied to internal memos as well as regulatory submissions. Second, the accept/reject standard: every AI edit is deliberately approved or refused by a named person, never bulk-accepted. Third, a documented list of artefacts that stay human-authored, owned by the quality lead and reviewed twice a year.
Source: erichgrunewald.substack.com