Football goalposts sliding across a chart of decade-old AI predictions, showing shifting AI goalposts

In September 2016, a Hacker News commenter named Houshalter argued that “seed AI” was a long way off, because it would first require an AI as intelligent as a human AI researcher. We are nowhere near that point, he wrote. Nine years later, 54 people voted on whether that had happened. The votes split. Nobody could agree on where the line even was anymore, let alone whether we had crossed it.

That is the problem with shifting AI goalposts. Capability that felt impossible in 2016 became routine by 2023, and most of us quietly updated our assumptions without noticing. If you run quality or operations, that memory drift costs you real money: you reject automation ideas based on a mental model from three years ago. Here is how to build a written record of your own AI assumptions, and what to do with it.

The Capability You Called Impossible in 2016 Is Now a Line Item in Your Budget

The stoppels.ch goalposts project does something uncomfortably simple. It digs up old Hacker News AI comments stating flatly what machines would never do, then puts each one to a public vote: Yes, Not sure, No. Has this happened? Voting closed on a Friday afternoon in October, and the answers scattered.

That scatter is the tell. Nobody keeps a written record of what they believed AI could not do, so skepticism and hype get re-argued from zero every quarter. Last year’s “never” quietly becomes this year’s procurement item, and no one logs the change.

If you are deciding whether to automate inspection reporting this year, that amnesia has a price. The call gets made on a vague sense of what AI “can” do rather than on evidence you actually wrote down.

Screenshot of dated Hacker News comments beside modern AI output, showing shifting AI goalposts

What the Goalposts Vote Actually Measures

From sprawling comment thread to a single yes/no claim

The method is crude and that is its strength. Take an archived comment, strip out the hedging and the thread replies, and reduce it to one claim you can actually test. Then ask people to rule on it.

Houshalter’s 2016 comment gets compressed to a single line: “An AI is at least as intelligent as human AI researchers.” The original argument was conditional and careful. The reduced version is binary, and that is where it gets uncomfortable, because a conditional prediction can never be wrong while a binary one can.

Most AI assumptions inside manufacturing organisations sit in the first category. “We’ll look at vision inspection when the models get good enough.” Good enough for what, measured how, decided by whom? Until someone writes the threshold down, nobody can be held to it and nobody learns anything when it passes.

Why a 54-person vote is still useful signal

Be honest about the sample. Fifty-four voters, voting closed Friday 2 October at 4:00 PM, self-selected from a technical audience. This is a sentiment snapshot, not a benchmark suite, and it tells you nothing about model performance on any specific task. Treat it as a measurement of human agreement, not machine capability.

That is still worth something. Proper AI capability benchmarks measure what systems can do. This measures whether a roomful of informed professionals can even agree on what was claimed, which turns out to be the harder problem. When 54 people read one sentence and split three ways, the sentence was never a usable standard to begin with.

Apply that test to your own roadmap. If you handed your AI assumptions to five people on your leadership team and asked each one whether the condition had been met, how many would answer the same way? If the answer is fewer than all five, you do not have a decision criterion. You have a vibe with a deadline attached.

Met, Unmet, and Genuinely Contested: Three Buckets of AI Claims

Run through enough archived predictions and they sort themselves into three piles. Claims that quietly came true while everyone was arguing about something else. Claims that plainly did not happen and look worse every year. And claims that cannot be settled at all, because nobody ever agreed what counted as passing.

That third pile is the biggest, and it is why the vote offers a “Not sure” option at all. The option is not politeness. It is an admission that the claim was never written tightly enough to adjudicate.

Vague claims never resolve, they just get re-argued

Houshalter’s reduced claim is the clean example: “An AI is at least as intelligent as human AI researchers.” Intelligent by what measure? Matching which researcher, on which task, over what time horizon? You can hold any position on that sentence in 2025 and defend it for an hour.

Claims about understanding, reasoning, or real intelligence behave the same way. There is no finish line, so the argument restarts every time a new model ships. This is where goalposts move most freely, because moving them costs nothing when there was never a fixed marker in the ground.

Your own internal assumptions have the same defect. “AI can’t handle our complex inspection decisions” is not a prediction. It is a mood, and it will survive any amount of contrary evidence.

Testable claims age into obvious yes or obvious no

Claims with acceptance criteria resolve themselves. Transcribe shop floor audio at a stated word error rate. Read a handwritten batch record and match a field-level accuracy threshold. Classify surface defects against a labelled set at a specified false negative rate. Nobody votes “not sure” on those.

Write your assumptions the same way and they stop being opinions. State the task, the data, the metric, and the number you would need to see before you changed your mind.

Then the claim does the work for you. Twelve months later it either cleared the bar or it did not, and you spend zero meeting time re-litigating what you used to believe.

Three-column chart sorting AI claims into met, unmet, and contested shifting AI goalposts

Build Your Own Goalpost Ledger for AI on the Factory Floor

Do the same exercise internally, on your own plant. Get your operations and quality leads in a room and have them write down, today, the specific things they believe AI cannot do in your environment. Not opinions about AI in general. Tasks on your floor, with your data, at your line speed.

Four claims worth starting with: read handwritten shift logs from the night crew, draft a deviation report a QA lead signs without edits, classify defect images fast enough to keep up with the line, and answer an auditor’s question directly from your document stack. Some of those may already be solved. Write them down anyway. The point is the record, not the verdict.

Writing claims with pass/fail acceptance criteria

A claim that cannot fail is worthless. Houshalter’s 2016 comment was careful and conditional, which is exactly why 54 people still could not agree on it nine years later. Your internal claims need the opposite quality: brittle enough to break on a specific date.

  • The task: one named activity, scoped to one line or one document type.
  • The threshold: the accuracy, throughput, or edit rate that counts as passing.
  • The owner: one person who runs the test and reports the result.
  • The date: when the claim gets checked, written when the claim is made.

Thresholds should match what a competent human does today, not perfection. If your QA lead edits 20 percent of hand-drafted deviation reports, that is your bar. Set it higher and you are building an excuse, not a test.

The quarterly review that kills stale objections

Every quarter, pull up the ledger and run the tests. Mark each claim met, unmet, or untested. Untested counts as a failure of process, not evidence for either side.

The real value shows up in planning season. When someone says AI is not there yet, you ask which ledger entry they mean and when it was last checked. Debates that used to eat an hour end in two minutes. Objections now carry an expiry date, and stale ones stop steering your roadmap.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

The ROI of Knowing Exactly When a Capability Crossed the Line

A ledger of dated claims changes your detection speed. Teams that review their own assumptions on a schedule notice a capability flip within a quarter. Teams running on memory notice it when a competitor mentions it, or when a vendor demo makes them feel behind, which is usually a year and a half later.

The cost of a capability you noticed 18 months late

Put a number on it with your own figures. Take one reporting workflow, count the hours it consumes per month, multiply by loaded cost. That is what you pay every month you keep funding manual effort against a task the technology quietly started handling. Catch the shift before your next audit cycle and the automation is already in place when the auditor walks in. Catch it late and you have paid for the same work twice.

The opposite mistake costs more. Assume every claim has already flipped, deploy into a process where the technology genuinely is not ready, and you get a failed rollout plus a shop floor that will not touch the next one. Scepticism and over-enthusiasm fail for the identical reason: nobody wrote down what “ready” meant. Houshalter’s 2016 comment held up for years precisely because it named a condition. Most internal assumptions name nothing.

The discipline runs the same in both directions. Date every claim. Define the test that would settle it, in terms your QA lead would accept. Put the review on the calendar, quarterly, with an owner.

If you have one afternoon, spend it on a single claim. Pick the task your team is most confident AI cannot do yet, write the claim in one testable sentence, and define what evidence would make you change your mind. Then run the test against real production data, not a sample someone cleaned up for the demo. One afternoon gives you a dated entry and a method you can repeat. That is enough to stop re-arguing the same question every quarter.

Source: stoppels.ch

Leave a Reply