AgentDish directory
AI Research AI Tools
Accepted listings in this category.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#510
↑ +2
Prometheus
An autonomous research system that runs on a single workstation and aggressively checks its own claims with adversarial self-verification, replication, and calibration audits. |
AI Research / Autonomous Research Systems | 86 | ↑ +2 | 52 days ago | Details |
|
#1263
↓ -2
Socrates
Open-source multi-agent protocol for AI research agents. It pairs a tool-using Scientist with a question-only advisor that can only ask questions and approve plans, and the README includes quick-start setup plus notes on reproducing results on MLE-bench/Kaggle tasks. |
AI Research / Multi-agent systems | 82 | ↓ -2 | 67 days ago | Details |
|
A research post from Telem AI measuring cold-start latency across nine web-search APIs, with cache behavior, tail latency, and agent-query comparisons. |
AI Research / Benchmarking / Latency Research | 81 | ↑ +2 | 8 days ago | Details |
|
#1368
↑ +2
EuroMesh
A sourced model and short report exploring whether Europe could train a sovereign frontier AI model using public compute it already owns, with reproducible code, datasets, and a PDF report. |
AI Research / Analysis / Reports | 81 | ↑ +2 | 77 days ago | Details |
|
#1385
↓ -55
MarCognity-AI
An open-source research framework for structured LLM evaluation, claim verification, and source-grounded reflective reasoning. The repo describes modular components for retrieval, semantic scoring, skeptical claim checking, and benchmark-style epistemic assessment. |
AI Research / Evaluation / Verification Framework | 81 | ↓ -55 | 118 days ago | Details |
|
Google Research article describing Science One Framework, an autonomous research prototype focused on verifiable AI-generated research through evidence chains and claim verification. |
AI Research / Autonomous Research | 80 | ↓ -1 | 22 days ago | Details |
|
#1477
↑ +2
Two Eyes Arguing
An interactive AI experiment where two photorealistic eyes argue about everyday moral scenarios, grounded in the sociology of justification and symbolic boundaries. |
AI Research / Interactive Demo | 79 | ↑ +2 | 31 days ago | Details |
|
#1589
↑ +6
MiroThinker
MiroThinker is a science-focused AI research app that emphasizes prediction, verification, and evidence-backed answers. The page also points to a MiroMind app and suggests use cases across finance, medicine, and regulation. |
AI Research / Deep Research Agent | 78 | ↑ +6 | 80 days ago | Details |
|
arXiv paper describing AVA, a GenAI platform for policy and development research built on 4,000+ World Bank reports. The abstract highlights multilingual support, evidence-based synthesis, citation verifiability, and reasoned abstention when queries cannot be supported. |
AI Research / Trustworthy Generative AI | 78 | ↑ +6 | 95 days ago | Details |
|
#1613
↑ +6
Agora-1: The Multi-Agent World Model
Agora-1 is a multi-agent world model from Odyssey that simulates shared real-time environments for up to four participants, human or AI, with a focus on gaming, robotics, reinforcement learning, and foundation model research. |
AI Research / World Models | 78 | ↑ +6 | 104 days ago | Details |
|
Apple Machine Learning Research paper proposing LaDiR, a reasoning framework that combines a VAE-based latent space with latent diffusion to improve LLM text reasoning and iterative refinement. |
AI Research / LLM Reasoning | 78 | ↑ +5 | 117 days ago | Details |
|
#1642
→ 0
Maith
Maith is an open-source research workspace for exploring open mathematics with AI while keeping proof standards explicit and checkable. |
AI Research / Mathematics | 77 | → 0 | 41 days ago | Details |
|
#1714
↓ -1
Hyperagents
Research paper introducing hyperagents, a self-referential agent framework that combines a task agent and a meta agent into one editable program. The abstract describes a DGM-based system that improves both task performance and its own improvement process across domains. |
AI Research / Self-Improving Agents | 76 | ↓ -1 | 100 days ago | Details |
|
A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records. |
AI Research / LLM Evaluation & Analysis | 75 | → 0 | 97 days ago | Details |
|
#1741
↓ -1
Extrapolation Under an Exact Verifier
A research page describing a verified circle-packing result produced by a frozen language model with an external memory loop, with published trace, attempts archive, and repository code/data. |
AI Research / LLM Experiments | 74 | ↓ -1 | 18 hours ago | Details |
|
#1785
↓ -1
GPT Guesses Between 1 and 100
A GitHub research project that measures how gpt-4.1 responds when asked to pick a random number between 1 and 100, using 10,000 API calls and comparing the results to a uniform baseline. |
AI Research / Model Behavior Analysis | 74 | ↓ -1 | 98 days ago | Details |