AgentDish directory
agent-evaluation
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#341
↓ -4
Mirrors
Mirrors is a staging and replay environment for AI agents. It rebuilds the systems an agent calls, replays real sessions against prompts, tools, and model changes, and surfaces regressions before deployment. |
Developer Tools / AI Development Tools | 87 | ↓ -4 | 32 days ago | Details |
|
#385
↓ -4
commensa-audit
A local-first Python tool that audits Git history to quantify rework in AI-generated engineering work, including PR correction rates, churn clusters, superseded work, and line survival. |
Developer Tools / AI Analytics | 87 | ↓ -4 | 79 days ago | Details |
|
#498
↑ +2
AI Arcade
An interactive site showcasing short AI-built arcade games and a model bench that compares coding models across the same tiny game constraints. |
Developer Tools / Code Assistant | 86 | ↑ +2 | 49 days ago | Details |
|
#845
↓ -6
ClaySeal Arena
A prompt-injection capture-the-flag arena for testing how AI agents behave under attack. Users can play challenges, create guarded agents, add optional guards and sandboxed tools, and track scores on a leaderboard. |
Developer Tools / Code Assistant | 84 | ↓ -6 | 46 days ago | Details |
|
#1178
↓ -2
Measuring Autonomous AI Research
A research post from Prime Intellect comparing 153 autonomous runs across 18 frontier models on a nanoGPT optimizer speedrun. It presents the setup, harness, results, and discussion around autonomous AI research performance. |
Research / AI Research Evaluation | 82 | ↓ -2 | 15 days ago | Details |
|
#1541
↑ +6
dsh-plugin-llm-verifier
A DeepSeek Harness plugin that uses an LLM to verify and rank candidate outputs with select, compare, track, and rollout tools. |
Developer Tools / AI DevTools | 78 | ↑ +6 | 11 days ago | Details |
|
#1771
↓ -1
Evaluate Your Agentic Tooling
A blog post describing an evaluation harness for comparing agentic coding tools and prompts across realistic SWE tasks, with token-cost results and model-specific behavior notes. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 78 days ago | Details |