AgentDish directory

agent-evaluation

Accepted listings with this tag.

Listing Category Score Trend Checked
#341 ↓ -4
Mirrors

Mirrors is a staging and replay environment for AI agents. It rebuilds the systems an agent calls, replays real sessions against prompts, tools, and model changes, and surfaces regressions before deployment.

Developer Tools / AI Development Tools 87 ↓ -4 32 days ago Details
#385 ↓ -4
commensa-audit

A local-first Python tool that audits Git history to quantify rework in AI-generated engineering work, including PR correction rates, churn clusters, superseded work, and line survival.

Developer Tools / AI Analytics 87 ↓ -4 79 days ago Details
#498 ↑ +2
AI Arcade

An interactive site showcasing short AI-built arcade games and a model bench that compares coding models across the same tiny game constraints.

Developer Tools / Code Assistant 86 ↑ +2 49 days ago Details
#845 ↓ -6
ClaySeal Arena

A prompt-injection capture-the-flag arena for testing how AI agents behave under attack. Users can play challenges, create guarded agents, add optional guards and sandboxed tools, and track scores on a leaderboard.

Developer Tools / Code Assistant 84 ↓ -6 46 days ago Details

A research post from Prime Intellect comparing 153 autonomous runs across 18 frontier models on a nanoGPT optimizer speedrun. It presents the setup, harness, results, and discussion around autonomous AI research performance.

Research / AI Research Evaluation 82 ↓ -2 15 days ago Details

A DeepSeek Harness plugin that uses an LLM to verify and rank candidate outputs with select, compare, track, and rollout tools.

Developer Tools / AI DevTools 78 ↑ +6 11 days ago Details

A blog post describing an evaluation harness for comparing agentic coding tools and prompts across realistic SWE tasks, with token-cost results and model-specific behavior notes.

Developer Tools / Code Assistant 74 ↓ -1 78 days ago Details