AgentDish directory

dpo

Accepted listings with this tag.

Listing Category Score Trend Checked
#792 ↓ -6
nano-llm-posttraining

Minimal LLM post-training experiments on a single 8GB GPU, covering SFT, DPO, and GRPO with a focus on forgetting, seed variance, and RL behavior.

AI Development / LLM Training & Fine-tuning 84 ↓ -6 30 days ago Details