AgentDish directory
evaluation
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records. |
AI Research / LLM Evaluation & Analysis | 75 | → 0 | 98 days ago | Details |
|
#1752
↓ -1
Jekyll-Hyde
A Hermes plugin that uses adversarial LLM clones to confront sandbagging and reward-hacking behavior during agent sessions. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 23 days ago | Details |
|
#1803
↑ +1
WifeBench
A playful benchmark dashboard that ranks LLMs based on one person's 10-question scoring process. |
Writing / Copywriting | 73 | ↑ +1 | 58 days ago | Details |