AgentDish directory
evaluation
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#3
↓ -1
Traccia
Traccia is an AI agent observability and governance platform with OpenTelemetry-native tracing, runtime policy enforcement, prompt registry, evals, cost attribution, and compliance evidence export. |
Developer Tools / Code Assistant | 92 | ↓ -1 | 10 days ago | Details |
|
#62
↓ -2
ForecastOps
Local-first observability and evaluation layer for production forecasts. Captures forecasts, validates them, computes horizon-aware metrics, and serves a read-only local UI for comparing runs. |
Developer Tools / MLOps / Observability | 90 | ↓ -2 | 80 days ago | Details |
|
#118
→ 0
ReactBench
ReactBench is a benchmark for evaluating coding agents on realistic React work, with published scores, cost comparisons, and example tasks focused on production-grade frontend issues. |
Developer Tools / Code Assistant | 89 | → 0 | 47 days ago | Details |
|
#129
→ 0
Mirrors
Mirrors turns production traces into an isolated, runnable copy of an AI agent’s environment so teams can replay workflows, reproduce bugs, and test changes safely before shipping. |
Developer Tools / AI Agent Testing | 89 | → 0 | 60 days ago | Details |
|
#165
↓ -3
Understudy
Understudy is a scenario-driven testing framework for AI agents. It simulates multi-turn users, records traces of messages and tool calls, and evaluates agent behavior with deterministic checks, optional LLM judges, and reports. |
Developer Tools / AI Testing | 88 | ↓ -3 | 4 days ago | Details |
|
#302
↓ -3
CAD-Bench
An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition. |
Research / Knowledge Work | 88 | ↓ -3 | 115 days ago | Details |
|
#307
↓ -3
agent-skills-eval
A TypeScript CLI and SDK for testing whether Agent Skills improve model outputs by running with-skill vs baseline evaluations and generating reports. |
Developer Tools / AI Evaluation | 88 | ↓ -3 | 117 days ago | Details |
|
#613
↑ +923
Alignment Whack-a-Mole
A research code repository for studying how fine-tuning can trigger verbatim recall of copyrighted books in large language models. It includes preprocessing, fine-tuning, generation, and memorization-evaluation scripts, with setup notes and example data. |
Research / Copywriting | 86 | ↑ +923 | 119 days ago | Details |
|
#730
↓ -6
Agentic Determinism Index
Open-source harness for measuring how consistently hosted LLM APIs return the same output across repeated runs and over time, with raw transcripts, scoring, and a static leaderboard generator. |
Developer Tools / Code Assistant | 84 | ↓ -6 | 5 hours ago | Details |
|
#765
↓ -6
Agent Memory Leaderboard
Public benchmark for comparing agent memory systems, with separate textual and coding tracks, submission flow, documentation, and API guidance. |
Developer Tools / Benchmark | 84 | ↓ -6 | 19 days ago | Details |
|
#866
↓ -6
Agent Infra
A curated GitHub resource list for production AI agent infrastructure, covering runtimes, sandboxes, tool protocols, security, observability, and evaluation. |
Developer Tools / AI Infrastructure | 84 | ↓ -6 | 57 days ago | Details |
|
#942
↓ -6
DeepSWE
DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks. The page shows a leaderboard, methodology overview, task examples, and a full blog explaining the benchmark design and results. |
Developer Tools / AI Benchmarking | 84 | ↓ -6 | 96 days ago | Details |
|
A GitHub repository that publishes a week of controlled runs comparing seven Google models on the same software task, with transcripts, telemetry, screenshots, and scored benchmark artifacts. |
AI Development / Benchmarking | 83 | ↓ -3 | 8 days ago | Details |
|
#1220
↓ -2
Ingot
Local-first change control for agent skills, with evidence-gated promotion, SkillOpt optimization, and MCP-based serving of skills. |
Developer Tools / AI Development | 82 | ↓ -2 | 41 days ago | Details |
|
#1271
↓ -2
clawmark
A local Rust CLI for A/B testing two CLAUDE.md files against a fixed SWE-bench Lite smoke set, with doctor, run, and report commands. |
Developer Tools / AI Benchmarking | 82 | ↓ -2 | 75 days ago | Details |
|
#1342
↑ +2
iFixAi
Open-source auditing tool for AI agents that checks whether an agent is doing the job it is supposed to do. The repo shows a guided CLI, plugin-based workflows for Claude Code and Codex, and output in JSON, Markdown, and terminal scorecards. |
Developer Tools / Code Assistant | 81 | ↑ +2 | 20 days ago | Details |
|
#1352
↑ +2
BOUND
BOUND is a deterministic control harness for AI agents. It sits between agent execution and the next decision, using observable evidence to choose ACCEPT, RETRY, REPLAN, or ROLLBACK. |
Developer Tool / AI Agent Framework / Control Harness | 81 | ↑ +2 | 45 days ago | Details |
|
#1479
↑ +2
Marker
Marker is an AI testing and verification platform for voice agents, with a focus on simulation, transcripts, judgments, and versioned evidence across the workflow. |
Developer Tools / AI Testing & Evaluation | 79 | ↑ +2 | 33 days ago | Details |
|
#1509
↑ +2
ArXiv Scholar
A zero-budget research search engine and RAG pipeline for 5,600 arXiv papers, built with free Colab processing, Qdrant free tier, and hybrid dense+sparse retrieval. |
AI / Search / RAG / Research search | 79 | ↑ +2 | 77 days ago | Details |
|
#1546
↑ +6
aakit
Aakit is a developer tool for measuring and tracking assumptions made by coding agents, linking each assumption to supporting code and flagging which ones fail. The repo includes a clear explanation of the idea, prior art, and early experiment results. |
Developer Tools / Code Assistant | 78 | ↑ +6 | 20 days ago | Details |
|
A research page on how censorship and behavior transfer during model distillation, with published models, data, evaluation code, and a benchmark called LineageEval. |
Research / AI Safety / Model Behavior | 78 | ↑ +6 | 32 days ago | Details |
|
A GitHub example that audits LangChain’s RAG quickstart with retrieval-quality metrics, flags off-topic and out-of-distribution queries, and surfaces ranking and calibration issues with charts and results files. |
Developer Tool / RAG Evaluation | 78 | ↑ +4 | 118 days ago | Details |
|
#1637
→ 0
agent-memory-leaderboard
Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems. |
Developer Tool / AI Evaluation / Benchmarking | 77 | → 0 | 30 days ago | Details |
|
#1725
→ 0
Clusy
Clusy is an agent-native notebook platform for ML and data science work in the cloud. The page says it can source data, inspect it, choose architecture and compute, and run end-to-end workflows, with a demo showing a finetuning task and follow-up work queued while the notebook runs. |
Research / Knowledge Work | 75 | → 0 | 62 days ago | Details |