Research / LLM Evaluation

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks

A research preprint and dataset on how five frontier LLMs disagree when judging 1,000 real-world fact-check claims, with accompanying corpus, raw results, and code repository.

Clear26/30
Useful24/30
Specific15/20
Complete11/20
Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks screenshot

Why it was accepted

The page clearly presents a focused AI research project with strong evidence: it studies frontier LLM judgment on factual claims, includes quantified results, and links to the underlying corpus and repository. That makes it useful for readers looking for LLM evaluation work, datasets, or reproducible research.

Weakness

The crawl shows the preprint, dataset files, and repository link, but not the code layout, installation steps, or example outputs from the GitHub project itself. It also does not explain what is inside the CSV beyond the fact that it contains raw results.

Review status

19 days ago #1683 ↓ -1

Last evaluated 19 days ago. Current rank #1683. Down 1 spot in the rankings.

Score history

76

Related listings

Primus AI Researcher – Free screenshot

Research / Knowledge Work

Primus is an autonomous AI researcher that hypothesizes, reads papers, writes code, runs experiments on compute, and drafts research papers. The page shows example tasks, published-paper claims, waitlist access, and positioning for ML research workflows.

Below the Fold — A New York Times X-Ray Dashboard screenshot

Research / Data Visualization

An interactive dashboard that analyzes New York Times coverage since 2000 using the NYT Archive API, with views for reporters, beats, sections, subjects, geography, obituaries, and corrections.

CAD-Bench screenshot
#302 CAD-Bench
88

Research / Knowledge Work

An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition.

Benchmarking Inference Engines on Agentic Workloads screenshot

Research / Knowledge Work

A research article from Applied Compute on how agentic, tool-using workloads differ from traditional LLM benchmarks, with production observations, workload profiles, and an open-source harness for replaying traces.