AgentDish directory

Research AI Tools

Accepted listings in this category.

Listing Category Score Trend Checked

Primus is an autonomous AI researcher that hypothesizes, reads papers, writes code, runs experiments on compute, and drafts research papers. The page shows example tasks, published-paper claims, waitlist access, and positioning for ML research workflows.

Research / Knowledge Work 90 ↓ -2 21 days ago Details

An interactive dashboard that analyzes New York Times coverage since 2000 using the NYT Archive API, with views for reporters, beats, sections, subjects, geography, obituaries, and corrections.

Research / Data Visualization 89 ↑ +454 118 days ago Details
#302 ↓ -3
CAD-Bench

An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition.

Research / Knowledge Work 88 ↓ -3 114 days ago Details

A research article from Applied Compute on how agentic, tool-using workloads differ from traditional LLM benchmarks, with production observations, workload profiles, and an open-source harness for replaying traces.

Research / Knowledge Work 87 ↓ -107 117 days ago Details
#442 ↑ +2
ThoughtDAG

An open-source, local-first canvas for editing LLM context as a graph. It lets users branch, prune, merge, and inspect the exact context sent to a model, with a desktop app, web demo, and support for Ollama and OpenAI-compatible endpoints.

Research / Knowledge Work 86 ↑ +2 16 days ago Details
#515 ↑ +2
UnderstandDocs

UnderstandDocs is a document analysis tool that summarizes pasted text or uploaded files, flags risks, extracts important dates, and simplifies dense language. The page shows a working analysis form, supported file types, a privacy/no-storage claim, and an example output for a tenancy agreement.

Research / Knowledge Work 86 ↑ +2 54 days ago Details

Anthropic research article on Claude interpretability, presenting evidence for a J-space and a global-workspace-like mechanism in language models.

Research / Interpretability 86 ↑ +2 55 days ago Details
#530 ↑ +2
ZUSE Automat Agent

A deterministic Python project for empirical law discovery in elementary cellular automata, with simulation, discovery, reproducibility guides, and published preprints.

Research / Scientific Discovery 86 ↑ +2 63 days ago Details
#613 ↑ +923
Alignment Whack-a-Mole

A research code repository for studying how fine-tuning can trigger verbatim recall of copyrighted books in large language models. It includes preprocessing, fine-tuning, generation, and memorization-evaluation scripts, with setup notes and example data.

Research / Copywriting 86 ↑ +923 118 days ago Details
#746 ↓ -6
Valovest

Valovest is an AI stock sentiment analysis tool that synthesizes analyst reports, earnings calls, news, and social sentiment into a short investment brief. The page shows timeframe filters, stock examples, a weekly featured stock, and outputs like overall sentiment, bull vs. bear arguments, and key themes.

Research / Knowledge Work 84 ↓ -6 7 days ago Details
#856 ↓ -6
Clusy

Clusy is an agent-native notebook platform for ML and data science that lets users describe a goal in plain language and have the system source data, set up experiments, run notebook cells, and return editable results in the cloud.

Research / Knowledge Work 84 ↓ -6 53 days ago Details

arXiv paper describing QUEST, an open family of deep research agents from 2B to 35B parameters, plus a synthetic-task training recipe and released models, data, and scripts.

Research / AI Agents 83 ↓ -3 97 days ago Details
#1126 ↓ -3
wwwatch

A daily AI intelligence journal for builders, covering notable model, tooling, and release updates in a short sourced digest.

Research / Knowledge Work 83 ↓ -3 101 days ago Details
#1130 ↓ -3
Physics AI

Physics AI is a physics homework and study tool that solves problems from photos or typed prompts, with step-by-step explanations, tutor mode, and visual breakdowns for diagrams and vectors.

Research / Knowledge Work 83 ↓ -3 102 days ago Details

A research report on the current MCP ecosystem, with live crawl numbers, verification rates, category breakdowns, and examples of both strong and weak MCP-positive sites.

Research / AI research 83 ↑ +174 118 days ago Details

A research post from Prime Intellect comparing 153 autonomous runs across 18 frontier models on a nanoGPT optimizer speedrun. It presents the setup, harness, results, and discussion around autonomous AI research performance.

Research / AI Research Evaluation 82 ↓ -2 15 days ago Details
#1265 ↓ -2
The Cascade Graph

An interactive, cited knowledge graph that maps physical, geopolitical, and economic constraints across the global economy, with node evidence pages, feedback loops, and an accessible text index.

Research / Knowledge Work 82 ↓ -2 69 days ago Details
#1280 ↓ -2
BigTech AI News

Chrome extension that tracks major AI companies, pulls in AI news and research, and generates daily summaries with Gemini, including language-aware summaries and article deep dives.

Research / Knowledge Work 82 ↓ -2 82 days ago Details
#1320 ↑ +213
ShadowBrokers

AI-powered trade signal product for retail traders that turns financial news into ranked trade plans with entries, stops, targets, and tracked accuracy.

Research / Knowledge Work 82 ↑ +213 117 days ago Details

A research page comparing 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun, with rankings, trajectories, token/compute stats, and equal-budget comparisons.

Research / Knowledge Work 81 ↑ +2 8 days ago Details

A research article from Plicara Labs analyzing 1.9 million GitHub agent skills and how often they include code, with breakdowns by language, mention-to-code gaps, and writing language.

Research / Data Analysis 80 ↓ -1 44 hours ago Details
#1391 ↓ -1
Canvas Chat

Canvas Chat is a tree-based LLM chat interface for branching, comparing, and organizing conversations on an infinite canvas using your own API keys.

Research / Knowledge Work 80 ↓ -1 4 days ago Details

A position paper arguing that AI alignment methods can be repurposed for censorship and manipulation, with examples across pre-training, post-training, and inference-time controls.

Research / Knowledge Work 79 ↑ +2 54 days ago Details

A research page on how censorship and behavior transfer during model distillation, with published models, data, evaluation code, and a benchmark called LineageEval.

Research / AI Safety / Model Behavior 78 ↑ +6 31 days ago Details