AgentDish directory

LLM evaluation

Accepted listings with this tag.

Listing Category Score Trend Checked
#732 ↓ -6
Preseason

Open-source benchmark that measures which developer tools LLMs recommend when asked to build real web apps. It runs fixed prompts against a panel of models, parses the tool recommendations, and publishes rankings and head-to-head comparisons.

Developer Tools / AI Benchmarking 84 ↓ -6 22 hours ago Details
#744 ↓ -6
Crilio

Crilio is a Python CLI and CI gate for testing AI prompts and catching regressions before they ship. It versions prompt tests in crilio.yaml, runs model calls, uses a judge model to score responses, and can block pull requests in GitHub Actions when checks fail.

Developer Tools / AI Testing & Evaluation 84 ↓ -6 6 days ago Details
#768 ↓ -6
LLM-Tests

An open-source benchmark for testing whether LLMs can follow long arithmetic derivations without using tools. It compares models on recall-proof inputs, logs digit accuracy, and reports results across multiple endpoints and model families.

Developer Tools / Code Assistant 84 ↓ -6 19 days ago Details
#1385 ↓ -55
MarCognity-AI

An open-source research framework for structured LLM evaluation, claim verification, and source-grounded reflective reasoning. The repo describes modular components for retrieval, semantic scoring, skeptical claim checking, and benchmark-style epistemic assessment.

AI Research / Evaluation / Verification Framework 81 ↓ -55 118 days ago Details
#1505 ↑ +2
AdvertBench

AdvertBench is a web app for ranking AI-generated image ad sets with Elo voting. The page shows head-to-head ad comparisons, a leaderboard, and a sample prompt for generating ads, making the product purpose easy to understand.

Developer Tools / Code Assistant 79 ↑ +2 71 days ago Details

A research preprint and dataset on how five frontier LLMs disagree when judging 1,000 real-world fact-check claims, with accompanying corpus, raw results, and code repository.

Research / LLM Evaluation 76 ↓ -1 19 days ago Details

Pull request adding TY25 support for moonshotai/kimi-k3 in TaxCalcBench through OpenRouter. The snapshot shows benchmark wiring, PDF handling details, saved outputs and evaluation reports, and regenerated charts/results for the full TY25 set.

AI Developer Tool / Benchmarking / evaluation 76 ↓ -1 40 days ago Details
#1729 → 0
Giskard

Giskard presents an AI red-teaming and continuous evaluation platform focused on finding hallucinations, security issues, and agent vulnerabilities. The page explains how it compares tools for agent-level testing, regression checks, and guardrail workflows.

Security / AI Red Teaming 75 → 0 68 days ago Details