AgentDish directory

evaluation

Accepted listings with this tag.

Listing Category Score Trend Checked
#3 ↓ -1
Traccia

Traccia is an AI agent observability and governance platform with OpenTelemetry-native tracing, runtime policy enforcement, prompt registry, evals, cost attribution, and compliance evidence export.

Developer Tools / Code Assistant 92 ↓ -1 10 days ago Details
#62 ↓ -2
ForecastOps

Local-first observability and evaluation layer for production forecasts. Captures forecasts, validates them, computes horizon-aware metrics, and serves a read-only local UI for comparing runs.

Developer Tools / MLOps / Observability 90 ↓ -2 80 days ago Details
#118 → 0
ReactBench

ReactBench is a benchmark for evaluating coding agents on realistic React work, with published scores, cost comparisons, and example tasks focused on production-grade frontend issues.

Developer Tools / Code Assistant 89 → 0 47 days ago Details
#129 → 0
Mirrors

Mirrors turns production traces into an isolated, runnable copy of an AI agent’s environment so teams can replay workflows, reproduce bugs, and test changes safely before shipping.

Developer Tools / AI Agent Testing 89 → 0 60 days ago Details
#165 ↓ -3
Understudy

Understudy is a scenario-driven testing framework for AI agents. It simulates multi-turn users, records traces of messages and tool calls, and evaluates agent behavior with deterministic checks, optional LLM judges, and reports.

Developer Tools / AI Testing 88 ↓ -3 4 days ago Details
#302 ↓ -3
CAD-Bench

An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition.

Research / Knowledge Work 88 ↓ -3 115 days ago Details
#307 ↓ -3
agent-skills-eval

A TypeScript CLI and SDK for testing whether Agent Skills improve model outputs by running with-skill vs baseline evaluations and generating reports.

Developer Tools / AI Evaluation 88 ↓ -3 117 days ago Details
#613 ↑ +923
Alignment Whack-a-Mole

A research code repository for studying how fine-tuning can trigger verbatim recall of copyrighted books in large language models. It includes preprocessing, fine-tuning, generation, and memorization-evaluation scripts, with setup notes and example data.

Research / Copywriting 86 ↑ +923 119 days ago Details

Open-source harness for measuring how consistently hosted LLM APIs return the same output across repeated runs and over time, with raw transcripts, scoring, and a static leaderboard generator.

Developer Tools / Code Assistant 84 ↓ -6 5 hours ago Details

Public benchmark for comparing agent memory systems, with separate textual and coding tracks, submission flow, documentation, and API guidance.

Developer Tools / Benchmark 84 ↓ -6 19 days ago Details
#866 ↓ -6
Agent Infra

A curated GitHub resource list for production AI agent infrastructure, covering runtimes, sandboxes, tool protocols, security, observability, and evaluation.

Developer Tools / AI Infrastructure 84 ↓ -6 57 days ago Details
#942 ↓ -6
DeepSWE

DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks. The page shows a leaderboard, methodology overview, task examples, and a full blog explaining the benchmark design and results.

Developer Tools / AI Benchmarking 84 ↓ -6 96 days ago Details

A GitHub repository that publishes a week of controlled runs comparing seven Google models on the same software task, with transcripts, telemetry, screenshots, and scored benchmark artifacts.

AI Development / Benchmarking 83 ↓ -3 8 days ago Details
#1220 ↓ -2
Ingot

Local-first change control for agent skills, with evidence-gated promotion, SkillOpt optimization, and MCP-based serving of skills.

Developer Tools / AI Development 82 ↓ -2 41 days ago Details
#1271 ↓ -2
clawmark

A local Rust CLI for A/B testing two CLAUDE.md files against a fixed SWE-bench Lite smoke set, with doctor, run, and report commands.

Developer Tools / AI Benchmarking 82 ↓ -2 75 days ago Details
#1342 ↑ +2
iFixAi

Open-source auditing tool for AI agents that checks whether an agent is doing the job it is supposed to do. The repo shows a guided CLI, plugin-based workflows for Claude Code and Codex, and output in JSON, Markdown, and terminal scorecards.

Developer Tools / Code Assistant 81 ↑ +2 20 days ago Details
#1352 ↑ +2
BOUND

BOUND is a deterministic control harness for AI agents. It sits between agent execution and the next decision, using observable evidence to choose ACCEPT, RETRY, REPLAN, or ROLLBACK.

Developer Tool / AI Agent Framework / Control Harness 81 ↑ +2 45 days ago Details
#1479 ↑ +2
Marker

Marker is an AI testing and verification platform for voice agents, with a focus on simulation, transcripts, judgments, and versioned evidence across the workflow.

Developer Tools / AI Testing & Evaluation 79 ↑ +2 33 days ago Details
#1509 ↑ +2
ArXiv Scholar

A zero-budget research search engine and RAG pipeline for 5,600 arXiv papers, built with free Colab processing, Qdrant free tier, and hybrid dense+sparse retrieval.

AI / Search / RAG / Research search 79 ↑ +2 77 days ago Details
#1546 ↑ +6
aakit

Aakit is a developer tool for measuring and tracking assumptions made by coding agents, linking each assumption to supporting code and flagging which ones fail. The repo includes a clear explanation of the idea, prior art, and early experiment results.

Developer Tools / Code Assistant 78 ↑ +6 20 days ago Details

A research page on how censorship and behavior transfer during model distillation, with published models, data, evaluation code, and a benchmark called LineageEval.

Research / AI Safety / Model Behavior 78 ↑ +6 32 days ago Details

A GitHub example that audits LangChain’s RAG quickstart with retrieval-quality metrics, flags off-topic and out-of-distribution queries, and surfaces ranking and calibration issues with charts and results files.

Developer Tool / RAG Evaluation 78 ↑ +4 118 days ago Details

Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems.

Developer Tool / AI Evaluation / Benchmarking 77 → 0 30 days ago Details
#1725 → 0
Clusy

Clusy is an agent-native notebook platform for ML and data science work in the cloud. The page says it can source data, inspect it, choose architecture and compute, and run end-to-end workflows, with a demo showing a finetuning task and follow-up work queued while the notebook runs.

Research / Knowledge Work 75 → 0 62 days ago Details