AgentDish directory
benchmark
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#11
↓ -3
gemma4.c
A pure C runtime for Gemma 4 E2B CPU inference, with export, benchmarking, and numerical validation tools included. |
Developer Tools / Machine Learning / Inference Runtime | 91 | ↓ -3 | 3 days ago | Details |
|
#118
→ 0
ReactBench
ReactBench is a benchmark for evaluating coding agents on realistic React work, with published scores, cost comparisons, and example tasks focused on production-grade frontend issues. |
Developer Tools / Code Assistant | 89 | → 0 | 47 days ago | Details |
|
#168
↓ -3
Robot Football League (RFL)
An open robot football league where frontier AI models and open-source club code manage simulated Unitree G1 humanoids in daily matches, with live streams, match logs, and a public engine for teams to join. |
Developer Tools / Code Assistant | 88 | ↓ -3 | 6 days ago | Details |
|
#227
↓ -3
Genesys
Open-source causal-graph memory for AI agents with MCP support, scoring, active forgetting, and benchmark results published in the repo. |
AI Development / Memory / Agent Infrastructure | 88 | ↓ -3 | 41 days ago | Details |
|
#302
↓ -3
CAD-Bench
An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition. |
Research / Knowledge Work | 88 | ↓ -3 | 115 days ago | Details |
|
#429
↑ +2
Modded-NanoGPT
Open-source repository for a NanoGPT training speedrun, focused on getting a 124M model to target loss on 8x H100 GPUs in under 90 seconds. The README includes run commands, Docker instructions, and a detailed list of model and systems optimizations. |
Developer Tools / Machine Learning / Training Optimization | 86 | ↑ +2 | 9 days ago | Details |
|
#517
↑ +2
The Banana Test
A visual benchmark that asks AI coding agents to generate a single-file Three.js animation of a banana plant’s full life cycle, then compares the live results side by side. |
AI Tools / Benchmarking | 86 | ↑ +2 | 56 days ago | Details |
|
#730
↓ -6
Agentic Determinism Index
Open-source harness for measuring how consistently hosted LLM APIs return the same output across repeated runs and over time, with raw transcripts, scoring, and a static leaderboard generator. |
Developer Tools / Code Assistant | 84 | ↓ -6 | 5 hours ago | Details |
|
#732
↓ -6
Preseason
Open-source benchmark that measures which developer tools LLMs recommend when asked to build real web apps. It runs fixed prompts against a panel of models, parses the tool recommendations, and publishes rankings and head-to-head comparisons. |
Developer Tools / AI Benchmarking | 84 | ↓ -6 | 29 hours ago | Details |
|
#765
↓ -6
Agent Memory Leaderboard
Public benchmark for comparing agent memory systems, with separate textual and coding tracks, submission flow, documentation, and API guidance. |
Developer Tools / Benchmark | 84 | ↓ -6 | 19 days ago | Details |
|
#768
↓ -6
LLM-Tests
An open-source benchmark for testing whether LLMs can follow long arithmetic derivations without using tools. It compares models on recall-proof inputs, logs digit accuracy, and reports results across multiple endpoints and model families. |
Developer Tools / Code Assistant | 84 | ↓ -6 | 20 days ago | Details |
|
#942
↓ -6
DeepSWE
DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks. The page shows a leaderboard, methodology overview, task examples, and a full blog explaining the benchmark design and results. |
Developer Tools / AI Benchmarking | 84 | ↓ -6 | 96 days ago | Details |
|
A GitHub repository that publishes a week of controlled runs comparing seven Google models on the same software task, with transcripts, telemetry, screenshots, and scored benchmark artifacts. |
AI Development / Benchmarking | 83 | ↓ -3 | 8 days ago | Details |
|
A black-box benchmark report on how AI-generated tests detect functional bugs in live APIs across 20 scenarios and 7 systems. |
Developer Tools / Code Assistant | 83 | ↓ -3 | 89 days ago | Details |
|
#1349
↑ +2
Vestige – Silent Rotation benchmark
A GitHub repo section documenting a benchmark for Vestige, a local-first Rust MCP memory layer for multi-agent coding fleets. The page explains the benchmark setup, what is being measured, how to reproduce the main claim, and where to find transcripts, evidence, and harness code. |
Developer Tools / AI / ML Infrastructure | 81 | ↑ +2 | 42 days ago | Details |
|
A workbench report comparing MiniMax M3 and GLM 5.2 on autonomous coding tasks, with scored results, latency and cost data, task-type breakdowns, and examples of where each model performed better. |
Developer Tools / Code Assistant | 81 | ↑ +2 | 73 days ago | Details |
|
#1413
↓ -1
AI World Bakeoff
A showcase of 27 explorable Three.js worlds built autonomously by 9 coding models from the same brief. |
Design / Creative Tools | 80 | ↓ -1 | 42 days ago | Details |
|
A Topos case study showing how a guided refactor of a synthetic healthcare claims engine reduced token usage, wall time, and estimated cost for later Gemini feature sessions. |
Developer Tools / Code Assistant | 78 | ↑ +6 | 67 days ago | Details |
|
#1637
→ 0
agent-memory-leaderboard
Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems. |
Developer Tool / AI Evaluation / Benchmarking | 77 | → 0 | 30 days ago | Details |
|
An in-depth Medium post from IronBee about how a verification loop affects AI coding agents, using Web-Bench and comparing DeepSeek with Claude Opus on a real web app task. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 56 days ago | Details |
|
#1803
↑ +1
WifeBench
A playful benchmark dashboard that ranks LLMs based on one person's 10-question scoring process. |
Writing / Copywriting | 73 | ↑ +1 | 58 days ago | Details |