What Tablif read on nerfWatch()'s site
nerfWatch() describes itself as: “Daily AI benchmarks and community votes on model quality”
- Stage: Live, no waitlist
What it is
Calls itself “AI benchmarks”.
- Calls itself: AI benchmarks
What it costs
No monthly plan.
- Monthly: no
“API calls for every test are paid out of pocket.”
“Contact made by @sergehere”
Who it sells to
Sells to developers.
- Sells to: developers
“We test what developers get through the API, not chat apps like ChatGPT or Claude.ai.”
How it runs
A API. Hosted in the cloud. Offers an API. Open source. AI is the core of the product.
- Kind: api
- Runs where: cloud / hosted
- API: yes
- Open source: yes
- AI role: core to the product
“…nerfWatch() tests AI models daily through their APIs”
“…nerfWatch() tests AI models through their public APIs.”
What it works with
Integrates with OpenRouter. Names ChatGPT, Claude Opus 5.5 and 4 more on its pages.
- Integrates with: OpenRouter
- Names: ChatGPT · Claude Opus 5.5 · Claude.ai · +3
“ChatGPT”
“Claude Opus 5.5”
What it offers as proof
Live, no waitlist. Claims “100 questions”.
- Number claimed: 100 questions
“Every model card has two buttons for that, Nerfed and Feels fine.”
“100 questions”
Sources
Tablif read these pages, last on 9 Oct 2026. Everything above is what the product says about itself; Tablif does not verify every claim.
Similar products in Tablif Town
- Compute:Arena: Community submitted benchmarks for Local AI
- NerfTracker: We show you when models get nerfed
- The AI Leaderboard: Compare the world’s leading AI models with benchmarks.
- Unbenchmark: Introducing Unbenchmark
- Benchmark Registry: One place for AI benchmark results.
- Feedback Bench by Enterpret: AI agents, ranked weekly by what users actually say
- modelsentiment: Which AI model does Reddit actually like?
- Elevist: AI daily execution system for high performers