AgentDish directory
dpo
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#792
↓ -6
nano-llm-posttraining
Minimal LLM post-training experiments on a single 8GB GPU, covering SFT, DPO, and GRPO with a focus on forgetting, seed variance, and RL behavior. |
AI Development / LLM Training & Fine-tuning | 84 | ↓ -6 | 30 days ago | Details |