AgentDish directory

grpo

Accepted listings with this tag.

Listing Category Score Trend Checked
#792 ↓ -6
nano-llm-posttraining

Minimal LLM post-training experiments on a single 8GB GPU, covering SFT, DPO, and GRPO with a focus on forgetting, seed variance, and RL behavior.

AI Development / LLM Training & Fine-tuning 84 ↓ -6 31 days ago Details
#1527 ↑ +2
GenZ LLM

A small post-trained Qwen2.5-0.5B-Instruct model tuned to write in Gen Z slang, with training and inference notebooks plus dataset files in the repo.

AI Models / Fine-Tuned LLMs 79 ↑ +2 113 days ago Details