AgentDish directory
ai-evaluation
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#1043
↓ -3
ARF — a record format for evaluation runs
ARF defines a JSON record format for AI evaluation runs with canonicalization and reproducible SHA-256 digests, so scores can be traced back to the underlying question, model, inputs, and judgment. |
Developer Tool / AI Evaluation / Data Format | 83 | ↓ -3 | 25 days ago | Details |
|
#1488
↑ +2
Lagotto Meter
Lagotto Meter is a beta web tool that checks how an AI agent perceives a website versus what the site claims about itself. It positions itself as a semantic truth/readability analyzer with a reproducible analysis method and public tools. |
Developer Tool / AI Evaluation | 79 | ↑ +2 | 42 days ago | Details |
|
#1610
↑ +6
LLM INQUISITOR
A GitHub repository that proposes a practical methodology for evaluating how AI systems behave during real work, with quick-start, practitioner, and methodology guides included. |
Developer Tools / AI Evaluation | 78 | ↑ +6 | 104 days ago | Details |
|
#1674
↑ +118
Agent Eval
A GitHub repo for evaluating agentic AI pipeline systems, with guidance for defining metrics, building eval cases, running repeatable tests, and tracking regressions. |
Developer Tools / Copywriting | 77 | ↑ +118 | 118 days ago | Details |