AgentDish directory

LLM evals

Accepted listings with this tag.

Listing Category Score Trend Checked

A research page comparing 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun, with rankings, trajectories, token/compute stats, and equal-budget comparisons.

Research / Knowledge Work 81 ↑ +2 8 days ago Details