AgentDish directory

benchmarking

Accepted listings with this tag.

Listing Category Score Trend Checked
#15 ↓ -3
SigMap

SigMap is a deterministic grounding layer for AI code work. It generates a signature-and-evidence map, helps pick relevant files, validates coverage, and judges whether an AI answer is grounded in the repo.

Developer Tools / AI Code Assistance 91 ↓ -3 57 days ago Details
#35 ↓ -2
BoundaryBench

BoundaryBench is an open-source benchmark for coding agents running under hardened sandbox policies. It compares harnesses like Claude Code, Codex, Terminus 2, and Grok on Terminal-Bench tasks and includes quickstart, policy controls, and result export tools.

Developer Tools / Code Assistant 90 ↓ -2 25 days ago Details
#77 ↓ -3
trycua/cua

Open-source infrastructure for computer-use agents, with sandboxes, SDKs, benchmarks, and desktop automation tooling for macOS, Linux, Windows, and Android. The repo also includes Cua Driver, CuaBot, Cua-Bench, and Lume for VM management.

Developer Tools / AI Agent Infrastructure 90 ↓ -3 118 days ago Details

오픈소스 AI 인프라 산정 도구로, 보유 GPU나 목표 워크로드를 입력해 실행 가능한 모델, 필요 장비, 배치안, TCO 비교까지 계산합니다. LLM, RAG, VLM, 아바타, 음성 워크로드를 다루며 한국어/영어 UI와 여러 추천 경로를 제공합니다.

Developer Tools / Code Assistant 88 ↓ -3 7 days ago Details
#185 ↓ -3
Is AI Dumber Today?

A community-tracked index of how users feel AI model quality is changing, updated hourly from public feedback and direct submissions.

AI Analytics / Model Tracking / Benchmarking 88 ↓ -3 17 days ago Details
#292 ↓ -3
LLMRequirements.com

An interactive guide for choosing local AI hardware and matching open-weights LLMs to specific builds. It offers budget-based build recommendations, a hardware picker, model-to-hardware compatibility, and a state-of-the-local-AI snapshot.

Developer Tools / AI / ML Infrastructure 88 ↓ -3 99 days ago Details
#307 ↓ -3
agent-skills-eval

A TypeScript CLI and SDK for testing whether Agent Skills improve model outputs by running with-skill vs baseline evaluations and generating reports.

Developer Tools / AI Evaluation 88 ↓ -3 116 days ago Details
#409 ↓ -4
SQLite-Columnar

A loadable SQLite extension that adds column-oriented storage and analytics for fast local OLAP-style queries, with benchmark data and build instructions.

Developer Tools / Databases & Storage 87 ↓ -4 110 days ago Details

A research article from Applied Compute on how agentic, tool-using workloads differ from traditional LLM benchmarks, with production observations, workload profiles, and an open-source harness for replaying traces.

Research / Knowledge Work 87 ↓ -107 118 days ago Details

A research repo for reducing hallucinations in LLM-generated code using semantic triangulation, with setup, benchmarking, experimentation, and reproducibility instructions.

Developer Tool / Code Quality 86 ↑ +2 22 days ago Details
#513 ↑ +2
Atelier

Open-source runtime for coding agents that sits underneath Claude Code to reduce tool calls, shorten context, and track savings. The page includes installation steps, a local savings check, and benchmark results comparing cost and performance.

Developer Tools / AI Development 86 ↑ +2 53 days ago Details
#517 ↑ +2
The Banana Test

A visual benchmark that asks AI coding agents to generate a single-file Three.js animation of a banana plant’s full life cycle, then compares the live results side by side.

AI Tools / Benchmarking 86 ↑ +2 55 days ago Details
#893 ↓ -6
Caplets

Caplets is a developer tool that wraps MCP servers into smaller capability-based surfaces for coding agents. The site explains the workflow, setup, example capabilities like OSV, GitHub, and Sourcegraph, and includes benchmark results plus docs links.

Developer Tools / Code Assistant 84 ↓ -6 68 days ago Details
#1008 ↓ -4
YourMemory

A persistent memory layer for AI agents, built as a standard MCP server with local setup, dashboard, and benchmark claims against other memory tools.

Developer Tools / AI Memory / MCP 84 ↓ -4 118 days ago Details

arXiv paper describing QUEST, an open family of deep research agents from 2B to 35B parameters, plus a synthetic-task training recipe and released models, data, and scripts.

Research / AI Agents 83 ↓ -3 97 days ago Details
#1187 ↓ -2
ExpertCache

Experimental page-aware Metal runtime and reproducibility harness for running oversized sparse mixture-of-experts models on Apple Silicon. The repo documents GPT-OSS 120B support, performance results on 64 GiB and 16 GiB machines, and the runtime/harness structure.

Developer Tools / AI / ML Infrastructure 82 ↓ -2 22 days ago Details
#1271 ↓ -2
clawmark

A local Rust CLI for A/B testing two CLAUDE.md files against a fixed SWE-bench Lite smoke set, with doctor, run, and report commands.

Developer Tools / AI Benchmarking 82 ↓ -2 74 days ago Details

A research page comparing 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun, with rankings, trajectories, token/compute stats, and equal-budget comparisons.

Research / Knowledge Work 81 ↑ +2 8 days ago Details
#1362 ↑ +2
PSI KV Governor

Reference implementation for using Linux Pressure Stall Information to trim an LLM KV cache under memory pressure. The repo includes requirements, basic usage commands, a simulator, a llama.cpp runner, and benchmark scripts with example results.

Developer Tools / AI Infrastructure 81 ↑ +2 64 days ago Details
#1461 ↑ +2
Agent Lightning v1.0.1

A GitHub release for Microsoft’s Agent Lightning Skill, which helps coding agents optimize other AI agents using benchmarks and iterative improvements across prompts, tools, workflows, models, and reasoning settings.

Developer Tools / AI Agents / Agent Training 79 ↑ +2 6 days ago Details
#1504 ↑ +2
HoprLabs

A Python CLI and research toolkit for simulating AI training math before model training. It estimates memory, training time, token budget, config risks, benchmark speed, and reliability, with optional native Rust and C backends.

Developer Tool / AI Research Toolkit 79 ↑ +2 66 days ago Details

An Apple Silicon–optimized inference build of Bonsai 1.7B with custom Metal kernels, benchmark results, quick-start instructions, and a bundled OpenAI-compatible server.

Developer Tools / Code Assistant 79 ↓ -208 117 days ago Details

A blog post about verifiable RAG that benchmarks open-source NLI verifiers against Claude on RAGTruth and describes a Python library for sentence-level citation and claim verification.

AI / RAG / Verification & Hallucination Detection 78 ↑ +6 90 days ago Details

A blog post from Augment Code comparing its coding agent, Auggie, against Claude Code on Opus 4.7. It presents benchmark results, token usage, cost comparisons, and an explanation of the Context Engine and Prism router.

Developer Tools / AI Coding Assistant 77 → 0 105 days ago Details