AgentDish directory

evaluation

Accepted listings with this tag.

Listing Category Score Trend Checked

A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records.

AI Research / LLM Evaluation & Analysis 75 → 0 98 days ago Details
#1752 ↓ -1
Jekyll-Hyde

A Hermes plugin that uses adversarial LLM clones to confront sandbagging and reward-hacking behavior during agent sessions.

Developer Tools / Code Assistant 74 ↓ -1 23 days ago Details
#1803 ↑ +1
WifeBench

A playful benchmark dashboard that ranks LLMs based on one person's 10-question scoring process.

Writing / Copywriting 73 ↑ +1 58 days ago Details