AgentDish directory

benchmarking

Accepted listings with this tag.

Listing Category Score Trend Checked

A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records.

AI Research / LLM Evaluation & Analysis 75 → 0 97 days ago Details

A technical blog post from the Cua repository describing a process-scoped Metal capability shim for macOS VMs that improves llama.cpp performance on Apple Silicon, with benchmark results and reproduction details.

Developer Tools / AI Infrastructure 74 ↓ -1 20 days ago Details

A research article describing an empirical runtime-monitoring approach for multi-turn LLM agents, backed by 3,175 runs across four benchmarks. It also points to an open-source implementation, state-harness, with Rust/Python support, LangGraph and CrewAI adapters, CLI tooling, and OpenTelemetry export.

Developer Tools / Code Assistant 74 ↓ -1 28 days ago Details

A blog post describing an evaluation harness for comparing agentic coding tools and prompts across realistic SWE tasks, with token-cost results and model-specific behavior notes.

Developer Tools / Code Assistant 74 ↓ -1 78 days ago Details
#1774 ↓ -1
LEVI

LEVI is a harness-first evolutionary framework for code and prompt optimization. It focuses on reducing LLM cost with diversity-preserving search, role-aware model routing, and a proxy benchmark, and presents comparative results against several existing systems.

Developer Tools / Code Assistant 74 ↓ -1 84 days ago Details

A Superconductor blog post showing how background coding agents were used to reproduce, diagnose, and fix a Rails memory leak using derailed_benchmarks, with a reusable Agent Skill workflow included.

Developer Tools / Code Assistant 74 ↓ -1 95 days ago Details

A Zenodo preprint reporting a 12,160-trial black-box evaluation of GPT-5.4 outputs under two prompt conditions, with detailed token-ceiling sweeps, control/ablation trials, and hash-chained verification.

Research / AI Model Evaluation 72 ↑ +1 26 days ago Details