AgentDish directory

post-training

Accepted listings with this tag.

Listing Category Score Trend Checked
#792 ↓ -6
nano-llm-posttraining

Minimal LLM post-training experiments on a single 8GB GPU, covering SFT, DPO, and GRPO with a focus on forgetting, seed variance, and RL behavior.

AI Development / LLM Training & Fine-tuning 84 ↓ -6 30 days ago Details

arXiv paper on distilling multi-agent debate into a single LLM with a two-stage fine-tuning pipeline. The abstract reports lower token use, comparable or better benchmark performance, and an analysis of agent-specific activation subspaces, with code linked from the page.

Research / AI/LLM Reasoning 74 ↓ -1 87 days ago Details