Courtroom gavel beside a screen showing AI-generated evidence of a reconstructed victim statement

An Arizona appeals court just threw out a 10-year sentence because of a video. Gabriel Paul Horcasitas was convicted of shooting Christopher Pelkey at a red light in 2021. Before sentencing, Pelkey’s family played an AI-generated video of him speaking, built from old recordings, photos and footage. The trial judge was moved by it. The appeals court was not, ruling that the clip “crossed that line” because it presented “a depiction of the victim and his thoughts created from the imaginings of the victim’s sister” rather than documenting an event.

That distinction, documenting what happened versus reconstructing what someone imagines happened, is the same test your AI-generated evidence will face in an audit, a deviation investigation, or a customer claim. Here is where the line falls, and how to stay on the right side of it.

A Judge Loved the AI Video. The Appeals Court Threw It Out Anyway.

Judge Todd Lang was not skeptical in the room. He was grateful.

“I loved that AI, thank you for that. As angry as you are, as justifiably angry as the family is, I heard the forgiveness. I feel that that was genuine.”

Genuine is the word that should worry you. The output felt true, so nobody in the courtroom stopped to ask what it actually was. An appeals court did, and the sentence came apart.

That gap is the whole problem with AI-generated evidence, and it shows up in operations the same way. A polished AI summary of a deviation investigation reads well, so it gets signed. Eighteen months later an auditor asks which parts were observed and which parts the model filled in. Persuasive is not the same as defensible, and nobody flags the difference while it’s happening.

Courtroom monitor displaying an AI-generated evidence video as a judge and attorneys watch

What Actually Happened in the Horcasitas Resentencing Ruling

A jury convicted Gabriel Paul Horcasitas, 55, of killing Christopher Pelkey, 37, in a 2021 confrontation at an Arizona red light. That verdict came from witnesses, physical evidence and the usual machinery of a criminal trial. The AI video played no part in it.

Pelkey’s sister, Stacey Wales, assembled the recreation afterwards, using voice recordings, videos and pictures of her brother. The synthetic version of Pelkey addressed the courtroom directly before sentencing, speaking about forgiveness. Defence lawyer Kristen Reller filed the appeal arguing the judge should never have allowed it.

The conviction stands; only the sentencing stage was tainted

This is the detail most coverage blurs. The Arizona Court of Appeals did not acquit anyone or reopen the question of guilt. It ordered a resentencing, nothing more. Horcasitas remains a convicted killer.

The ruling is narrow on purpose. One piece of material entered one stage of the process and contaminated that stage, so the court rewound to that point and no further. The ten-year term is gone, the finding of fact behind it is not.

Apply the same logic to your operations. When an AI-generated record fails scrutiny, you rarely lose the underlying event. You lose the decision that was built on top of it, and you redo the work that produced that decision. That is still expensive, just expensive in a specific and predictable way.

Why nobody objected at the time

Everyone in that courtroom knew the video was synthetic. Wales said openly that the family had built it. There was no deception, no hidden provenance, no attempt to pass off a reconstruction as a recording. It still got thrown out.

Disclosure was not the problem. The problem was that nobody in the room asked what category of thing they were looking at, because it was emotionally coherent and arrived at a moment when everyone wanted something like it to exist. Judge Lang accepted it on the spot.

Your reviewers do the same thing with a clean, confident AI summary. Nobody objects to output that reads correctly.

The Test the Court Applied: Documentation vs. Reconstruction

The ruling turns on one distinction, and it is portable to any records process you run. Either a document captures something that occurred, or it assembles a plausible version of something that might have. The Arizona Court of Appeals put it in a single line.

Rather than document an event or recording a particular moment, the AI video presents a depiction of the victim and his thoughts created from the imaginings of the victim’s sister.

Run that test against an AI-drafted deviation report, a root cause summary, or a shift handover written from sensor logs and operator notes. If the output asserts a cause, an intent, or a sequence that no input actually established, it is a reconstruction. Useful, maybe. A record, no.

Authentic inputs, invented output: how the gap opens

Every input in the Pelkey video was real. Voice recordings, photographs, footage, all genuinely his. The court still rejected the result, because authentic source material does not make the assembled output a recording of anything.

The same mechanic breaks quality records. Your CAPA inputs are real: batch data, operator statements, maintenance logs. The model stitches them into a narrative with a stated root cause, and that cause came from pattern-matching, not evidence. Nobody flagged it, because every component it drew on was legitimate.

The gap opens at the point of assembly, not the point of input. That is where your controls have to sit.

Why ‘it felt genuine’ is not a standard

Fluency is the failure mode. A well-structured AI summary reads like a competent engineer wrote it, which is precisely why reviewers stop interrogating it. Polish gets mistaken for provenance.

Auditors and regulators apply the same test the appeals court did. They ask what the record is derived from and whether that derivation can be traced. “The output was accurate and the team agreed with it” answers neither question.

Build the documentation-versus-reconstruction distinction into your review step explicitly. Ask of every AI-generated record: which claim here came from a source, and which came from the model?

Split-screen comparison of a documented crash photo and AI-generated evidence reconstructing the same scene

Where This Bites in Quality, Compliance and Operations

Your teams are already doing this. AI summaries of audit findings, CAPA narratives, incident write-ups, supplier assessments, training completion records. Each one is a document that a notified body, an insurer, an FDA investigator or opposing counsel may read years from now, with no memory of how it was produced.

The failure mode is not a hallucinated number. It is a well-written paragraph describing a root cause that nobody actually investigated, sitting in your quality system looking exactly like a paragraph that someone did investigate.

A four-item provenance checklist for AI-generated documents

Provenance is cheap to build at creation and nearly impossible to reconstruct later. Four controls cover most of the exposure.

  • Tag at generation: every AI-drafted record carries a machine-readable flag, model name and version, applied automatically rather than by the author remembering.
  • Link the inputs: store the source documents, logs or transcripts the output was built from, so anyone can check what the summary summarised.
  • Named human sign-off: one person, by name, attests that the content matches the underlying evidence before it leaves the building.
  • Separate drafts from the system of record: AI output lands in a staging area, not directly into your QMS entry.

The return here is not faster typing. It is avoided rework when an auditor asks how a finding was documented, and avoided findings when you can answer in thirty seconds instead of reopening twelve months of records.

The records that should never be AI-drafted

Anything that asserts a fact about the physical world that no one observed stays off-limits. Inspection results, test outcomes, deviation descriptions, witness statements from operators, calibration records. These document an event. An AI cannot document an event it was not present for, it can only produce a plausible account of one.

That is precisely the line the Arizona court drew: documenting a moment versus presenting something assembled “from the imaginings” of someone else. Apply it to your records and the sorting takes an afternoon.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

Provenance Becomes a Default Requirement, Not a Nice-to-Have

The question is shifting. For the past two years, auditors and reviewers have asked whether AI was involved in producing a record. That question is already obsolete, because the answer is almost always yes. The two questions replacing it are harder: what source material was this built from, and who put their name to it?

Arizona is an early marker, not an outlier. The appeals court did not object to the technology. It objected to a document whose inputs were someone’s recollection and imagination, presented with the authority of a recording. Any auditor can run that same test on a CAPA narrative, and some already do.

Organisations that can answer both questions in seconds will keep running AI at full speed. The rest will hit one bad finding and respond with a blanket internal ban, which is the worst possible outcome. You lose the productivity and keep none of the control, because people go back to doing it quietly in a browser tab.

What to put in place before your next audit cycle

Start with the metadata, not the policy document. Every AI-assisted record should carry three fields: the source inputs it drew on, the model and prompt used, and the named human who reviewed and approved it. If your QMS cannot hold those fields, that is a configuration problem you can fix this quarter.

  • Source lock: AI drafts cite only uploaded evidence, logs or interview notes. No general-knowledge filling of gaps.
  • Named sign-off: approval sits with a person who can explain the content without opening the prompt history.
  • Classification: split records into documented facts and interpretive analysis, and label them differently in the system.
  • Retrieval test: pick five AI-assisted records at random and reconstruct their provenance. If it takes more than a few minutes, you are not ready.

Build this now, while the consequence of failure is a finding in an internal audit report rather than a decision overturned on appeal.

Source: bbc.com

Leave a Reply