AgentDish directory

model benchmarking

Accepted listings with this tag.

Listing Category Score Trend Checked

A research post from Prime Intellect comparing 153 autonomous runs across 18 frontier models on a nanoGPT optimizer speedrun. It presents the setup, harness, results, and discussion around autonomous AI research performance.

Research / AI Research Evaluation 82 ↓ -2 15 days ago Details