AgentDish directory

reinforcement-learning

Accepted listings with this tag.

Listing Category Score Trend Checked
#792 ↓ -6
nano-llm-posttraining

Minimal LLM post-training experiments on a single 8GB GPU, covering SFT, DPO, and GRPO with a focus on forgetting, seed variance, and RL behavior.

AI Development / LLM Training & Fine-tuning 84 ↓ -6 30 days ago Details
#1527 ↑ +2
GenZ LLM

A small post-trained Qwen2.5-0.5B-Instruct model tuned to write in Gen Z slang, with training and inference notebooks plus dataset files in the repo.

AI Models / Fine-Tuned LLMs 79 ↑ +2 112 days ago Details
#1552 ↑ +6
EdotEnv (E.env)

Market-derived RL environments for training agents on quant trading and long-horizon planning under adversarial noise.

AI/ML / Reinforcement Learning 78 ↑ +6 26 days ago Details

Agora-1 is a multi-agent world model from Odyssey that simulates shared real-time environments for up to four participants, human or AI, with a focus on gaming, robotics, reinforcement learning, and foundation model research.

AI Research / World Models 78 ↑ +6 104 days ago Details

A blog post describing a small reinforcement-learning agent trained with PPO to play and beat a Pokelike/Pokerogue-style game, including the input representation, model architecture, and training approach.

Developer Tools / Code Assistant 74 ↓ -1 88 days ago Details

An educational article explaining world models, latent states, dynamics learning, and planning for agents, with examples from gridworld, Dreamer, and MuZero.

Writing / Copywriting 73 ↑ +1 106 days ago Details

arXiv paper on self-speculating LLM agents that predict their next tool call to hide tool latency and improve next-call Hit@1 while preserving task success.

Research / AI Paper 72 ↑ +1 33 days ago Details