AI Research / Evaluation / Verification Framework

MarCognity-AI

An open-source research framework for structured LLM evaluation, claim verification, and source-grounded reflective reasoning. The repo describes modular components for retrieval, semantic scoring, skeptical claim checking, and benchmark-style epistemic assessment.

Clear25/30
Useful23/30
Specific15/20
Complete18/20
MarCognity-AI screenshot

Why it was accepted

The page clearly presents an AI research project with a defined purpose: structured LLM evaluation and claim verification. The README gives enough substance for a public listing, including the framework’s modules, core capabilities, research motivation, and benchmark setup. It also shows concrete evidence of source retrieval, claim-level verification, and reproducible experimentation rather than a vague concept page.

Weakness

The snapshot cuts off before the full benchmark and usage details, so a visitor cannot see setup instructions, runnable examples, or results quality in full. It also does not fully show how the included modules are wired together in practice or what inputs and outputs look like end to end.

Review status

118 days ago #1385 ↓ -55

Last evaluated 118 days ago. Current rank #1385. Down 55 spots in the rankings.

Score history

78808379818281

Related listings

Prometheus screenshot
#510 Prometheus
86

AI Research / Autonomous Research Systems

An autonomous research system that runs on a single workstation and aggressively checks its own claims with adversarial self-verification, replication, and calibration audits.

Socrates screenshot
#1263 Socrates
82

AI Research / Multi-agent systems

Open-source multi-agent protocol for AI research agents. It pairs a tool-using Scientist with a question-only advisor that can only ask questions and approve plans, and the README includes quick-start setup plus notes on reproducing results on MLE-bench/Kaggle tasks.

The Web-Search Latency Your Agent Actually Pays screenshot

AI Research / Benchmarking / Latency Research

A research post from Telem AI measuring cold-start latency across nine web-search APIs, with cache behavior, tail latency, and agent-query comparisons.

EuroMesh screenshot
#1368 EuroMesh
81

AI Research / Analysis / Reports

A sourced model and short report exploring whether Europe could train a sovereign frontier AI model using public compute it already owns, with reproducible code, datasets, and a PDF report.