AgentDish directory

reproducibility

Accepted listings with this tag.

Listing Category Score Trend Checked

Open-source harness for measuring how consistently hosted LLM APIs return the same output across repeated runs and over time, with raw transcripts, scoring, and a static leaderboard generator.

Developer Tools / Code Assistant 84 ↓ -6 5 hours ago Details
#868 ↓ -6
Open Science

Open Science is an open-source AI workbench for scientists. It combines literature, code, figures, reports, and review into a local-first desktop workflow with model choice, reproducible artifacts, and scientific agent skills.

Developer Tools / AI Research Tools 84 ↓ -6 57 days ago Details

Open-source research skill that helps Claude, Codex, and OpenClaw run a more disciplined AI research workflow, from hypothesis and literature review through baselines, leakage checks, analysis, and paper writing.

AI Tools / AI Agents 84 ↓ -6 62 days ago Details

ARF defines a JSON record format for AI evaluation runs with canonicalization and reproducible SHA-256 digests, so scores can be traced back to the underlying question, model, inputs, and judgment.

Developer Tool / AI Evaluation / Data Format 83 ↓ -3 25 days ago Details