AgentDish directory

benchmark

Accepted listings with this tag.

Listing Category Score Trend Checked
#11 ↓ -3
gemma4.c

A pure C runtime for Gemma 4 E2B CPU inference, with export, benchmarking, and numerical validation tools included.

Developer Tools / Machine Learning / Inference Runtime 91 ↓ -3 3 days ago Details
#118 → 0
ReactBench

ReactBench is a benchmark for evaluating coding agents on realistic React work, with published scores, cost comparisons, and example tasks focused on production-grade frontend issues.

Developer Tools / Code Assistant 89 → 0 47 days ago Details

An open robot football league where frontier AI models and open-source club code manage simulated Unitree G1 humanoids in daily matches, with live streams, match logs, and a public engine for teams to join.

Developer Tools / Code Assistant 88 ↓ -3 6 days ago Details
#227 ↓ -3
Genesys

Open-source causal-graph memory for AI agents with MCP support, scoring, active forgetting, and benchmark results published in the repo.

AI Development / Memory / Agent Infrastructure 88 ↓ -3 41 days ago Details
#302 ↓ -3
CAD-Bench

An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition.

Research / Knowledge Work 88 ↓ -3 115 days ago Details
#429 ↑ +2
Modded-NanoGPT

Open-source repository for a NanoGPT training speedrun, focused on getting a 124M model to target loss on 8x H100 GPUs in under 90 seconds. The README includes run commands, Docker instructions, and a detailed list of model and systems optimizations.

Developer Tools / Machine Learning / Training Optimization 86 ↑ +2 9 days ago Details
#517 ↑ +2
The Banana Test

A visual benchmark that asks AI coding agents to generate a single-file Three.js animation of a banana plant’s full life cycle, then compares the live results side by side.

AI Tools / Benchmarking 86 ↑ +2 56 days ago Details

Open-source harness for measuring how consistently hosted LLM APIs return the same output across repeated runs and over time, with raw transcripts, scoring, and a static leaderboard generator.

Developer Tools / Code Assistant 84 ↓ -6 5 hours ago Details
#732 ↓ -6
Preseason

Open-source benchmark that measures which developer tools LLMs recommend when asked to build real web apps. It runs fixed prompts against a panel of models, parses the tool recommendations, and publishes rankings and head-to-head comparisons.

Developer Tools / AI Benchmarking 84 ↓ -6 29 hours ago Details

Public benchmark for comparing agent memory systems, with separate textual and coding tracks, submission flow, documentation, and API guidance.

Developer Tools / Benchmark 84 ↓ -6 19 days ago Details
#768 ↓ -6
LLM-Tests

An open-source benchmark for testing whether LLMs can follow long arithmetic derivations without using tools. It compares models on recall-proof inputs, logs digit accuracy, and reports results across multiple endpoints and model families.

Developer Tools / Code Assistant 84 ↓ -6 20 days ago Details
#942 ↓ -6
DeepSWE

DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks. The page shows a leaderboard, methodology overview, task examples, and a full blog explaining the benchmark design and results.

Developer Tools / AI Benchmarking 84 ↓ -6 96 days ago Details

A GitHub repository that publishes a week of controlled runs comparing seven Google models on the same software task, with transcripts, telemetry, screenshots, and scored benchmark artifacts.

AI Development / Benchmarking 83 ↓ -3 8 days ago Details

A black-box benchmark report on how AI-generated tests detect functional bugs in live APIs across 20 scenarios and 7 systems.

Developer Tools / Code Assistant 83 ↓ -3 89 days ago Details

A GitHub repo section documenting a benchmark for Vestige, a local-first Rust MCP memory layer for multi-agent coding fleets. The page explains the benchmark setup, what is being measured, how to reproduce the main claim, and where to find transcripts, evidence, and harness code.

Developer Tools / AI / ML Infrastructure 81 ↑ +2 42 days ago Details

A workbench report comparing MiniMax M3 and GLM 5.2 on autonomous coding tasks, with scored results, latency and cost data, task-type breakdowns, and examples of where each model performed better.

Developer Tools / Code Assistant 81 ↑ +2 73 days ago Details
#1413 ↓ -1
AI World Bakeoff

A showcase of 27 explorable Three.js worlds built autonomously by 9 coding models from the same brief.

Design / Creative Tools 80 ↓ -1 42 days ago Details

A Topos case study showing how a guided refactor of a synthetic healthcare claims engine reduced token usage, wall time, and estimated cost for later Gemini feature sessions.

Developer Tools / Code Assistant 78 ↑ +6 67 days ago Details

Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems.

Developer Tool / AI Evaluation / Benchmarking 77 → 0 30 days ago Details

An in-depth Medium post from IronBee about how a verification loop affects AI coding agents, using Web-Bench and comparing DeepSeek with Claude Opus on a real web app task.

Developer Tools / Code Assistant 74 ↓ -1 56 days ago Details
#1803 ↑ +1
WifeBench

A playful benchmark dashboard that ranks LLMs based on one person's 10-question scoring process.

Writing / Copywriting 73 ↑ +1 58 days ago Details