AgentDish directory
Research
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#1725
→ 0
Clusy
Clusy is an agent-native notebook platform for ML and data science work in the cloud. The page says it can source data, inspect it, choose architecture and compute, and run end-to-end workflows, with a demo showing a finetuning task and follow-up work queued while the notebook runs. |
Research / Knowledge Work | 75 | → 0 | 61 days ago | Details |
|
A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records. |
AI Research / LLM Evaluation & Analysis | 75 | → 0 | 97 days ago | Details |
|
#1741
↓ -1
Extrapolation Under an Exact Verifier
A research page describing a verified circle-packing result produced by a frozen language model with an external memory loop, with published trace, attempts archive, and repository code/data. |
AI Research / LLM Experiments | 74 | ↓ -1 | 21 hours ago | Details |
|
#1742
↓ -1
THROTTLE
An interactive simulation about how inference providers ration access to stronger LLMs when demand spikes. Users can compare the default industry rule with their own policy and explore who gets the strong model. |
AI Product / LLM Routing / Inference Management | 74 | ↓ -1 | 3 days ago | Details |
|
A Microsoft Research paper PDF on agentic coding and GitHub Copilot at production scale. The crawl confirms it is a substantive research document tied to an AI coding product and suitable for a public listing. |
Research / AI Research Paper | 74 | ↓ -1 | 22 days ago | Details |
|
arXiv paper introducing ScientistOne, an autonomous research system built around a chain-of-evidence framework to keep claims traceable and audit research outputs for verifiability. |
Research / AI Research Agents | 74 | ↓ -1 | 22 days ago | Details |
|
arXiv paper on a cache-merging method for multi-agent latent reasoning, framing KV-cache composition as a convergent replicated state with deterministic merging. |
Research / AI Research Paper | 74 | ↓ -1 | 59 days ago | Details |
|
Anthropic research report on how people use Claude Code in practice, based on a large privacy-preserving analysis of session data. It covers task types, division of labor between user and model, and how domain expertise affects outcomes. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 74 days ago | Details |
|
#1774
↓ -1
LEVI
LEVI is a harness-first evolutionary framework for code and prompt optimization. It focuses on reducing LLM cost with diversity-preserving search, role-aware model routing, and a proxy benchmark, and presents comparative results against several existing systems. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 84 days ago | Details |
|
arXiv paper on distilling multi-agent debate into a single LLM with a two-stage fine-tuning pipeline. The abstract reports lower token use, comparable or better benchmark performance, and an analysis of agent-specific activation subspaces, with code linked from the page. |
Research / AI/LLM Reasoning | 74 | ↓ -1 | 87 days ago | Details |
|
#1785
↓ -1
GPT Guesses Between 1 and 100
A GitHub research project that measures how gpt-4.1 responds when asked to pick a random number between 1 and 100, using 10,000 API calls and comparing the results to a uniform baseline. |
AI Research / Model Behavior Analysis | 74 | ↓ -1 | 98 days ago | Details |
|
#1792
↓ -1
The Cat Is Under Mayonnaise
An open-source experiment that adds a small zero-initialized overlay layer to a frozen GPT-2 so its behavior can be adjusted at inference time without retraining the base model. |
AI Developer Tool / Model Adaptation / Adapters | 74 | ↓ -1 | 117 days ago | Details |
|
MIT’s report on how generative AI is changing teaching, learning, assessment, and research training, with guiding principles and recommendations for instructors and the administration. |
Research / Knowledge Work | 73 | ↑ +1 | 4 days ago | Details |
|
arXiv paper on how ideas and goals can spread through multi-agent LLM systems, including experiments on propagation, susceptibility factors, and mitigation. |
Research / AI Safety | 72 | ↑ +1 | 10 days ago | Details |
|
A Zenodo preprint reporting a 12,160-trial black-box evaluation of GPT-5.4 outputs under two prompt conditions, with detailed token-ceiling sweeps, control/ablation trials, and hash-chained verification. |
Research / AI Model Evaluation | 72 | ↑ +1 | 26 days ago | Details |
|
arXiv paper on self-speculating LLM agents that predict their next tool call to hide tool latency and improve next-call Hit@1 while preserving task success. |
Research / AI Paper | 72 | ↑ +1 | 33 days ago | Details |
|
#1839
↑ +1
Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
An arXiv paper on reducing LLM inference cost for web automation by compiling browser tasks into a deterministic JSON workflow and executing them without repeated model calls. |
Research / Paper | 72 | ↑ +1 | 85 days ago | Details |
|
A GitHub-hosted research note about an agent-driven engineering workflow, comparing merged PR volume across Kungfu, Google AX, and OpenAI references. |
Writing / Copywriting | 71 | → 0 | 11 days ago | Details |
|
arXiv paper proposing MANTA, a framework that lets multi-agent communication structures adapt at inference time based on collaboration traces and task needs. |
Research / AI Paper | 71 | → 0 | 30 days ago | Details |
|
#1860
→ 0
Code as Agent Harness
A research page and preprint about using code as the runtime layer for agent systems, with a taxonomy of harness interfaces, harness mechanisms, and scaling patterns for multi-agent workflows. |
Developer Tools / Code Assistant | 71 | → 0 | 102 days ago | Details |