AgentDish directory
Research AI Tools
Accepted listings in this category.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#1558
↑ +6
BixRouter
A web app for branching AI conversations into a DAG-style chat tree, with model switching, response style controls, chat import/export, and example conversations. |
Research / Knowledge Work | 78 | ↑ +6 | 34 days ago | Details |
|
#1652
→ 0
Prompt Injection as Role Confusion
An ICML 2026 research project page arguing that prompt injection comes from how LLMs misread roles, with an extended writeup, examples, and links to the paper, code, arXiv, and BibTeX. |
Research / AI Safety | 77 | → 0 | 69 days ago | Details |
|
PaperProfit explains an AI-assisted stock evaluation approach that combines fundamentals, technical signals, and qualitative analysis from transcripts and SEC filings into a weighted score. |
Research / Knowledge Work | 77 | → 0 | 90 days ago | Details |
|
#1661
→ 0
CAPTCHAs can still detect AI agents
A research write-up on detecting AI agents through process differences in CAPTCHA and related cognitive tasks. It outlines the CogCAPTCHA30 approach, reports human-vs-model differences, and connects the findings to Roundtable’s Proof of Human product. |
Research / Knowledge Work | 77 | → 0 | 93 days ago | Details |
|
arXiv paper on a self-speculative decoding framework for speeding up reasoning LLM inference on edge hardware, with hardware co-design and reported speedups. |
Research / AI/ML Paper | 77 | → 0 | 94 days ago | Details |
|
#1677
↓ -1
Nanointerpret
An LLM interpretability playground for forcing activations, exploring learned concepts, and comparing unmodified vs intervened generations. |
Research / Knowledge Work | 76 | ↓ -1 | 4 days ago | Details |
|
A research preprint and dataset on how five frontier LLMs disagree when judging 1,000 real-world fact-check claims, with accompanying corpus, raw results, and code repository. |
Research / LLM Evaluation | 76 | ↓ -1 | 19 days ago | Details |
|
A public research report from Inkling that maps Apple-to-OpenAI employee moves using profile data, with month-by-month trends, function breakdowns, title shifts, and examples of the roles that moved. |
Research / Talent Intelligence | 76 | ↓ -1 | 41 days ago | Details |
|
#1710
↓ -1
PDF to MD Converter
AI-powered PDF to Markdown converter focused on preserving structure like headings, tables, images, and reading order for research, docs, and knowledge workflows. |
Research / Knowledge Work | 76 | ↓ -1 | 89 days ago | Details |
|
#1716
↓ -1
ITB Engine
A research repository and local web app for testing quantum gravity theory-space exclusions against encoded consistency constraints. It includes a CLI, a localhost interface, test coverage, and an LLM-powered research agent for running searches and report generation. |
Research / Scientific Computing | 76 | ↓ -1 | 112 days ago | Details |
|
#1725
→ 0
Clusy
Clusy is an agent-native notebook platform for ML and data science work in the cloud. The page says it can source data, inspect it, choose architecture and compute, and run end-to-end workflows, with a demo showing a finetuning task and follow-up work queued while the notebook runs. |
Research / Knowledge Work | 75 | → 0 | 61 days ago | Details |
|
A Microsoft Research paper PDF on agentic coding and GitHub Copilot at production scale. The crawl confirms it is a substantive research document tied to an AI coding product and suitable for a public listing. |
Research / AI Research Paper | 74 | ↓ -1 | 22 days ago | Details |
|
arXiv paper introducing ScientistOne, an autonomous research system built around a chain-of-evidence framework to keep claims traceable and audit research outputs for verifiability. |
Research / AI Research Agents | 74 | ↓ -1 | 22 days ago | Details |
|
#1758
↓ -1
SQLite Critical CVEs or LLM Slop?
JFrog Security Research investigates a batch of SQLite CVEs and argues many are AI-generated or unsupported by the source code and PoC testing. The post includes a comparison matrix, methodology, and detailed breakdown of several alleged vulnerabilities. |
Research / Security Research | 74 | ↓ -1 | 28 days ago | Details |
|
arXiv paper on a cache-merging method for multi-agent latent reasoning, framing KV-cache composition as a convergent replicated state with deterministic merging. |
Research / AI Research Paper | 74 | ↓ -1 | 59 days ago | Details |
|
arXiv paper on distilling multi-agent debate into a single LLM with a two-stage fine-tuning pipeline. The abstract reports lower token use, comparable or better benchmark performance, and an analysis of agent-specific activation subspaces, with code linked from the page. |
Research / AI/LLM Reasoning | 74 | ↓ -1 | 87 days ago | Details |
|
MIT’s report on how generative AI is changing teaching, learning, assessment, and research training, with guiding principles and recommendations for instructors and the administration. |
Research / Knowledge Work | 73 | ↑ +1 | 4 days ago | Details |
|
arXiv paper on how ideas and goals can spread through multi-agent LLM systems, including experiments on propagation, susceptibility factors, and mitigation. |
Research / AI Safety | 72 | ↑ +1 | 10 days ago | Details |
|
A Zenodo preprint reporting a 12,160-trial black-box evaluation of GPT-5.4 outputs under two prompt conditions, with detailed token-ceiling sweeps, control/ablation trials, and hash-chained verification. |
Research / AI Model Evaluation | 72 | ↑ +1 | 26 days ago | Details |
|
arXiv paper on self-speculating LLM agents that predict their next tool call to hide tool latency and improve next-call Hit@1 while preserving task success. |
Research / AI Paper | 72 | ↑ +1 | 33 days ago | Details |
|
Google Research paper describing a deployed multimodal defense system for detecting coordinated synthetic spam and abusive media at scale, using account relatedness, synthetic content classification, and LoRA-tuned LLMs. |
Research / AI Safety / Abuse Detection | 72 | ↑ +1 | 43 days ago | Details |
|
#1839
↑ +1
Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
An arXiv paper on reducing LLM inference cost for web automation by compiling browser tasks into a deterministic JSON workflow and executing them without repeated model calls. |
Research / Paper | 72 | ↑ +1 | 85 days ago | Details |
|
arXiv paper proposing MANTA, a framework that lets multi-agent communication structures adapt at inference time based on collaboration traces and task needs. |
Research / AI Paper | 71 | → 0 | 30 days ago | Details |