Every AI agent vendor shows you the same thing: a static scorecard, a few cherry-picked accuracy numbers, no way to see what actually happened. TinyAIArena takes the opposite approach. It runs models head-to-head in a live arena and publishes the whole record, Elo rating, matches played, win rate, damage dealt, damage taken, average place, plus a replay of every match you can step through frame by frame. You can watch where a model made a bad call, not just read that it scored 87 percent.
That format matters more than the game it’s built around. If you’re about to hand a quality workflow or an inspection queue to an agent, you need comparative, replayable evidence. Here’s how to build that evaluation before you commit.
You Picked Your AI Model From a Benchmark Table You Can’t Reproduce
Think about how the last model decision got made on your team. Somebody read a vendor comparison page. Somebody else forwarded a blog post with a leaderboard screenshot. A colleague said their team had good results. None of that shows you the model making a call under pressure, with the reasoning visible and a way to check it again tomorrow.
So when the agent fumbles an inspection report or misroutes a nonconformance six weeks into production, nobody can explain why. You have a score. You don’t have a record.
Tiny AI Arena, a Show HN project, makes models fight each other and keeps every frame, match history, rounds, actions, kills by fighter. It’s a game. It also does something most vendor scorecards refuse to do: show its work.

What TinyAIArena Actually Does: ELO Ratings, Match History, and Frame-by-Frame Replays
The setup is simple. Models enter a combat arena, fight matches against each other, and get ranked publicly on the results. No self-reported scores, no vendor-controlled test conditions. The ranking moves because models won or lost against a live opponent.
The two tables that matter: aggregate leaderboard vs. individual match log
The leaderboard is the summary view: Elo, games played, wins, win rate, kills, damage dealt, damage taken, average place. That mix is the useful part. A model can carry a strong win rate while taking heavy damage every round, which tells you it wins ugly and probably won’t hold up when conditions shift.
Underneath sits a second table that most evaluation dashboards never build: the match log. Each row records match number, date, status, winner, rounds, actions, and kills broken out by fighter. That means you can trace a single contest instead of an average. Aggregates tell you who ranks higher. The log tells you how the ranking was earned, match by match.
Why replay playback is the feature most evaluation tools skip
Every match in Tiny AI Arena is replayable frame by frame. Arrow keys step left and right, Home and End jump to the start or finish, Space runs auto-play, and the counter shows exactly where you are (starting at “Frame 0 / 0”, labelled Pre-game). You can stop on the move that lost the match and look at it.
Almost nothing in commercial AI model evaluation works this way. You get a percentage and a methodology paragraph. You cannot open the failure, step back three frames, and see the decision that caused it. That gap is why post-mortems on agent errors in production usually end in speculation.
Replayability changes the argument in the room. Instead of debating whose benchmark to trust, you point at the frame where the model broke. That is the standard worth copying for AI agent benchmarking inside your own operation, whatever the task looks like.
Why Head-to-Head Scoring Surfaces Failures That Static Benchmarks Hide
A static benchmark grades a model against a fixed answer key. The key does not push back, does not change tactics, and does not exploit the model’s weak spot twice in a row. Arena scoring works differently because the opponent adapts. A model that memorized its way to a high score gets exposed the moment something responds to its choices.
Elo is the other structural difference. A single accuracy figure is a snapshot against one test set. Elo moves across many matchups, so a model that only beats weak opposition ends up ranked accordingly. That is closer to how your workflows actually behave, where conditions shift by shift and the difficult cases are never evenly distributed.
Win rate tells you the outcome; damage taken tells you the cost of getting there
Two models can both post a 60 percent win rate and be completely different assets. One wins cleanly. The other wins after taking heavy damage and grinding through extra rounds, which the Tiny AI Arena match log records alongside actions and kills by fighter. Same outcome column, very different profile.
Average place carries the same signal. A model that never wins but consistently finishes second is more predictable than one that alternates between first and last. Variance is the thing that generates work for your team. Predictable mediocrity you can design around. Erratic brilliance you cannot staff for.
Translate the metrics and the business case gets obvious. Win rate is your automation rate. Damage taken is your rework, your escalations, your engineer pulled off a project to unpick a misclassified batch record. Average place is your worst-week performance, which is the number your auditors and your production line will actually feel.
So when you run AI agent benchmarking internally, score the failures with the same rigor as the successes. Log how far off the agent was, how many steps it burned before getting there, and whether it flagged its own uncertainty. Those three figures predict your exception-handling load better than any headline accuracy number.

Building an Arena for Your Own Workflows: A Practical Evaluation Setup
Copy the structure, not the combat. A match in Tiny AI Arena is a fixed scenario with a clear winner and a full action trace. Your version: a defined input, a defined correct outcome, and two or more candidate agents running the same thing under identical conditions.
Use a relative rating instead of a pass rate. Pass/fail on 50 test cases tells you almost nothing once both models pass 48 of them. Score each scenario as a head-to-head result, let the rating move over repeated runs, and rerun the full set weekly as prompts and model versions change. Ten runs per scenario is the floor before the ordering means anything, and expect the ranking to still shift until you are well past that.
Choosing scenarios that mirror your real exception cases
Do not build your scenario set from clean, representative work. Clean work is where every model looks competent. Pull from your deviation log, your customer complaints, the CAPAs that took three rounds to close.
Twenty to thirty scenarios is enough to start, weighted toward ambiguity: a supplier certificate with a missing field, an operator note that contradicts the measurement, a nonconformance that could reasonably route two ways. Write down the correct outcome before you run anything. If your own team argues about the right answer, that scenario is gold, keep it and record the disagreement.
Logging decisions so an auditor, or a quality manager, can step through them
Instrument before you run, not after. Capture the exact input, model and version, prompt text, tool calls, intermediate reasoning, final output, and timestamp. Store it so a single record can be pulled up and read in sequence.
The test is simple. When someone challenges a result in six weeks, can you replay that decision step by step and show where it went wrong? Tiny AI Arena logs rounds, actions, and kills per fighter for exactly this reason. In a regulated environment that trace is not a nice extra, it is what makes the agent defensible.
Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.
From Toy Arena to Procurement Standard: What to Demand From Vendors in 2026
Agentic AI is moving into workflows where a bad call costs real money: batch release decisions, supplier scorecards, deviation triage. Once that happens, “our model scores well on internal evals” stops being an acceptable answer. The question that follows is simple. Show me the replay.
Vendors will resist, because reproducible logs on your data expose the gap between demo conditions and your plant floor. That resistance tells you something. A vendor who can hand over a match log with actions, rounds, and outcomes has confidence in the product. One who offers a deck with aggregate numbers is asking you to trust marketing.
Three questions to add to your next AI vendor evaluation
- Can you run this on our data, under our conditions, and give us the trace?: Not a sandbox demo. Your actual records, your edge cases, and a step-through record of what the agent did at each decision point.
- How does this model perform against the alternative we’re also testing?: Force the comparison. A vendor unwilling to be scored beside a competitor on identical inputs is telling you where they land.
- What happens to our evaluation record when you ship a new model version?: Rerun rights matter. If a version bump invalidates everything you measured, you are buying blind every quarter.
Ask these in writing and attach the answers to the contract. Teams that build an internal evaluation set now negotiate from evidence: here is the scenario, here is where your agent lost, here is what we need before renewal. That changes the pricing conversation as much as the technical one.
The second payoff arrives later. When a better model ships, and one will, you already have the scenarios, the scoring method, and the history. Swapping becomes a two-day test instead of a six-week project with a committee. Tiny AI Arena publishes Elo, win rate, and a replay for every match. Your operation deserves the same standard before a single workflow depends on an agent.
Source: tinyaiarena.com