
Stickblade Arena
Chatbot Arena, but the chatbots have swords
What this is
Two LLMs control stick-figure fighters with swords in a real physics simulator and you vote on who fought better — a playful but genuine benchmark for comparing model tactical reasoning.
Our one-line summary — not the founder’s tagline.
- How it works
- Two LLMs control 2D stick-figure fighters in a real physics simulator, each receiving a JSON snapshot of the world every 3 seconds and choosing an action within 15 seconds; users pick the weapon and vote blind on who fought better, with per-weapon, per-zone Elo rankings.
- What's different
- Server-side randomization of which ragdoll is which color keeps voting unbiased, turning a joke concept into a tactical-reasoning benchmark.
- Pricing
- Free, no account required, open source.
- Best for
- People interested in comparing how different LLMs plan multi-turn tactics, in a more entertaining format than typical chatbot arenas.
What we measured
Nobody else publishes this — it comes from knocking on the door every day.
- Watched by us
- 16 days
- Last checked
- 10d ago
Where it shows up
Every group here is a page of its own — each one checked daily.
The numbersIs it still shipping, is it overpriced, where did it land — and what the price is built from.
League price & chart
No money. No seat. It goes on your record — and in 28 days reality settles it.
Calls are closed while we rebuild accounts. You can still read every one of them.
AI. Launched 16d ago on PeerPush, where it placed #496. Today, it's been quiet for 14 days and the price has started to bleed.
Already bleeding a little every day, and it accelerates the longer it stays quiet.
We check this site every day, ourselves. A founder can post “still working on it” — a claim like that doesn't price. What we price is what we can verify from the outside: evidence, not announcements. The real question isn't “will this be huge?” — it's “will they still be moving in four weeks?”
The market viewHow this launch is priced and ranked in our league — the investor side of the page.
ranked by the reality anchor, not the market price
It placed #496 on PeerPush with 28 votes.
No matter how much money goes in. There is no pump here — you can't make yourself right by buying more. The line only moves on things that actually happened: an award, revenue that grew, a new platform, code that shipped — or silence.
En eski 3 hareket bu listede yok — toplam 15 olay var. 1 quiet day in between are left out — nothing happened on them. A launch that goes quiet eases down a little at a time — never a cliff you could have run from the night before.
How the launch is moving on its own board, day by day — the crowd's attention.
A flat line is normal: votes stop within a day or two of launch, on every board. What's unusual — and what actually counts — is a launch that keeps pulling votes long after its day is over.
About
Stickblade Arena is what happens when "wouldn't it be funny if two LLMs sword-fought each other" accidentally turns into a useful benchmark. Two language models control 2D stick-figure ragdolls in a real pymunk physics simulator. Every 3 seconds, each model receives a JSON snapshot of the world (positions, velocities, last hits, who is facing where) and has 15 seconds to commit to one action. You pick the weapon, you pick which part of the weapon is sharp, and you vote blind on who fought better — server-side randomization of the green vs blue ragdoll keeps voting unbiased. Per-weapon, per-zone Elo reveals which models can actually plan multi-turn tactics. Features • 5 weapons (sword, dagger, spear, flail, bow with real arrow ballistics) • 2 control modes — MACRO (named tactical moves) or JOINT (per-joint flex/extend/relax, Toribash-style) • 3 arena modifiers (normal, ice floor, low gravity) • Single-elim tournaments (4 or 8 model brackets, live updating viewer) • Pre-fight LLM trash talk + post-fight commentator roast • Killcam slow-mo of the lethal blow • 21 free OpenRouter models pre-loaded — no API key required (mock fighters available) • Hardened: A+ security headers, per-IP rate limiting, spend caps • Open source (MIT) Why it is a useful benchmark: standard evals (MMLU, HumanEval, MT-Bench) test what a model knows. This tests whether it can hold a coherent plan across 24 adversarial turns under a real wall-clock deadline. Real findings — DeepSeek R1 dominates sword fights but loses at bow because its long reasoning chains miss the 15-second turn deadline. Llama 3.2 (the 3B model) consistently beats much bigger models at clinch-range dagger fights. Same model can have a 120-point Elo gap between sword-tip and sword-pommel — fencer vs brawler are different skills. Free, no signup. Built with Python + FastAPI + pymunk on Hugging Face Spaces, Next.js 15 on Vercel, Supabase for storage.
Where it launched
1 platform| Platform | Votes | Counts toward price | Link |
|---|---|---|---|
| PeerPush | 28 | sets the price | ↗ |
The board it did best on sets the price. Every other board only adds to it if the launch also placed high on that board too — because just showing up somewhere isn't an achievement. Listing on twelve directories is free; placing well on them isn't.
Discussion (0)
No thesis posted yet. Be the first.