Give yourself 60 seconds and 10 seconds per image, and try to separate real photos from AI-generated ones. That’s the entire premise of Reality Check, a browser game where a correct call earns +100 points and a wrong one costs you 100. Most people walk away with a lower score than they expected. It’s a fun way to lose an argument with yourself, and an uncomfortable preview of a control failure sitting inside your quality system right now.
If your team still eyeballs supplier photos, inspection evidence, or audit documentation to confirm they’re genuine, you’re relying on a verification method that no longer works. Below, we break down where visual judgement is still embedded in your processes, what the actual exposure looks like, and which controls replace it.
Your Inspection Photo Might Not Be a Photo
The game gives you one binary call per image: real photo or AI-generated. No forensic tools, no metadata, no second opinion. Just the same snap judgement your team makes every day, compressed into a scoreboard that tells you exactly how often you were wrong.
Now move that eye out of the browser. It’s the eye checking supplier condition photos before a shipment gets released. It’s the eye accepting non-conformance evidence from a plant three time zones away. It’s the eye reviewing incident documentation that ends up in an audit file. Same instinct, higher stakes, no score at the end to tell you it failed.
Image authenticity was never written into your quality system because it never needed to be. A photo was proof. That assumption expired quietly, and nobody updated the procedure.

What the Reality Check Game Actually Measures
Strip away the scoreboard and the game is a forced-choice detection test under time pressure with a penalty for guessing. That combination is rarer than it sounds. Most quizzes reward participation. This one punishes a wrong call as hard as it rewards a right one, which changes how you behave inside those ten seconds.
Time pressure and penalty scoring mirror a real sign-off decision
Ten seconds per image is roughly how long anyone looks at a supplier photo before deciding it’s fine. The clock never pauses. You can’t open the file in another tool, you can’t ask a colleague, and you can’t defer. Click 1 for real, 2 for AI, move on.
The scoring is where it gets honest. A timeout earns zero, so abstaining is cheap but useless, exactly like escalating every ambiguous photo to someone else. Guessing wrong costs 100 points against the 100 you’d have gained, so a coin flip is worth nothing over a round. Streak bonuses at 3×, 5× and 10× reward sustained accuracy rather than one lucky call, which is the property you actually want from an inspection control.
The post-round reveal is where the learning actually happens
When the 60 seconds expire, the game shows you your images, revealed. Every call you made, laid against the truth. That screen is worth more than the score, because the score only tells you that you were wrong and the reveal tells you where.
Patterns surface fast. Most people find they over-trust images with visible imperfection (grain, motion blur, bad lighting) and under-trust anything clean and well-composed. Both instincts are backwards now. Generative models produce mess on demand, and a sharp, evenly lit product shot is not evidence of a camera.
Run it with your quality team and compare reveal screens. You’ll get a cheap, specific map of which visual cues your people still trust and which ones have stopped working. That map is the starting point for deciding what your verification process should check instead.
Why Trained Eyes Still Get It Wrong
Detection failure runs in two directions, and most teams only worry about one of them. You can call a synthetic image real. You can also call a genuine one fake. Both errors put you in the same place: a sign-off you cannot defend.
The tells your team was trained to look for have been engineered out
Six fingers. Garbled text on signage. Suspiciously perfect symmetry. Every quality inspector who has read an article about spotting AI images is hunting for these, and generation models fixed them. The heuristics still feel reliable, which is worse than having no heuristic at all. Absence of a known tell now gets read as evidence of authenticity, and that is a logic error, not a perception error.
The reverse problem gets less attention and causes just as much damage. Real photos from a plant floor are compressed, badly lit, shot on a five-year-old phone under sodium vapour lamps. Sensor noise, JPEG banding, and motion blur all read as “something is off” to someone actively hunting for flaws. Your inspector rejects legitimate non-conformance evidence because it looks too rough to be true.
Push both error types together and accuracy converges on a coin flip. Reality Check makes this visible by charging you 100 points for a wrong call, the same amount a correct one earns. Guessing nets zero over enough rounds. That symmetry is the point: it strips out the illusion that being right half the time means you have a skill.
Here is the part that should bother an operations leader. The person making the call has no error rate. Not a bad one, none at all. They have never been scored, never been calibrated, and cannot tell you whether they are at 80 percent or 52 percent on image authenticity.
A control with an unmeasured error rate is not a control. It is an assumption wearing a signature block. You would never accept a gauge with no calibration record in your measurement system, and human visual judgement on provenance is exactly that gauge.

Where This Hits Manufacturing Operations First
Not every image in your business carries weight. A photo in a marketing deck is illustration. A photo attached to a batch release is evidence. The problem is that most quality systems never drew that line, so both get treated the same way: someone looks, someone approves, someone moves on.
The image-evidence inventory: which decisions actually depend on a picture
Start by listing every point where a picture triggers money moving or a record closing. In most manufacturing operations that list is longer than people expect.
- Supplier condition and packaging photos: released against a visual confirmation that pallets, seals, and labels are intact.
- Incoming goods inspection records: photographic proof attached to an accept or reject decision on a lot.
- Remote factory acceptance testing: a machine signed off on video and stills without anyone in the room.
- Warranty and returns claims: a payout decided on submitted damage photos.
- Subcontractor progress documentation: milestone invoices approved against site images.
- Customer-submitted defect reports: credits, replacements, and 8D investigations opened on a phone photo.
Rank each one by the value of the decision and the compliance exposure if the evidence turns out to be fabricated. That ranking is your actual risk register for image authenticity. Everything below the line can stay as it is. Everything above it needs a control that does not depend on someone’s instinct.
What a single fabricated image costs downstream
One altered packaging photo releases a damaged shipment into production. The line runs, the parts get built in, and you find out at final test or, worse, at the customer. The cost is not the image, it is the recall, the scrap, and the containment effort.
One fabricated damage photo gets a warranty claim paid that should have been rejected, and the same pattern repeats because it worked. One unverifiable inspection record turns into an audit finding on evidence integrity, which questions every other record in the file.
The practical move is unglamorous. Stop logging “a person reviewed it” as verification. Classify the decisions that carry real financial or compliance weight, then design a control for those specifically.
Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.
Building Verification That Does Not Rely on Eyeballs
Training your team to detect better is a losing strategy. Generation quality improves monthly; human perception does not. The fix is structural: stop asking reviewers to authenticate images and start controlling how images enter your system in the first place.
Shift verification from the reviewer to the capture point
An emailed JPEG carries no defensible history. A photo captured inside a controlled app carries a timestamp, GPS coordinates, device identifier, and an immutable upload record. That difference turns an opinion into a record, and it costs your supplier about thirty seconds more per photo.
Build the checks into intake rather than review. Reject loose files for evidence-grade decisions. Read C2PA content provenance manifests where they exist, flag images stripped of EXIF data, and apply chain-of-custody rules so evidence images are versioned and traceable from capture to audit file. Then define one escalation rule: when a picture is the sole basis for a high-value release, a credit, or a recall decision, it requires a second verification channel. Video, a live call, or a physical sample.
A 90-day sequence: inventory, harden, measure
Days 1 to 30: inventory. Pull your quality and operations team into a room and run the 60-second game first. Ten seconds per image, +100 for a correct call, minus 100 for a wrong one. Five minutes, and the argument about whether this is a real risk ends on its own. Then map every workflow where an image triggers money moving.
Days 31 to 60: harden the top three by financial exposure. Usually supplier acceptance, non-conformance evidence, and warranty claims. Days 61 to 90: measure rejection rates at intake and time added per submission.
The cost side is a workflow change measured in hours of configuration and a short supplier communication. The exposure side is a disputed claim, a recall built on false evidence, or an audit finding against your documentation integrity. Those land in six figures. The arithmetic is not difficult.
Source: slop-sense.labtoagi.com