AgentDish directory

LLM comparison

Accepted listings with this tag.

Listing Category Score Trend Checked

A GitHub repository that publishes a week of controlled runs comparing seven Google models on the same software task, with transcripts, telemetry, screenshots, and scored benchmark artifacts.

AI Development / Benchmarking 83 ↓ -3 7 days ago Details

A workbench report comparing MiniMax M3 and GLM 5.2 on autonomous coding tasks, with scored results, latency and cost data, task-type breakdowns, and examples of where each model performed better.

Developer Tools / Code Assistant 81 ↑ +2 72 days ago Details