Why it was accepted
The page is a real, substantive research artifact centered on AI model performance. It clearly explains the experiment, shows multiple comparison views, and exposes enough metrics and traces to be useful in a public listing.
Research / Knowledge Work
A research page comparing 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun, with rankings, trajectories, token/compute stats, and equal-budget comparisons.
The page is a real, substantive research artifact centered on AI model performance. It clearly explains the experiment, shows multiple comparison views, and exposes enough metrics and traces to be useful in a public listing.
The snapshot doesn’t explain the benchmark setup in depth, the scoring method behind the record-gap chart, or how to reproduce the runs from the page alone.
Last evaluated 8 days ago. Current rank #1338. Up 2 spots in the rankings.
Primus is an autonomous AI researcher that hypothesizes, reads papers, writes code, runs experiments on compute, and drafts research papers. The page shows example tasks, published-paper claims, waitlist access, and positioning for ML research workflows.
An interactive dashboard that analyzes New York Times coverage since 2000 using the NYT Archive API, with views for reporters, beats, sections, subjects, geography, obituaries, and corrections.
An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition.
A research article from Applied Compute on how agentic, tool-using workloads differ from traditional LLM benchmarks, with production observations, workload profiles, and an open-source harness for replaying traces.