ZeroShotMind

RL & Alignment

Reinforcement Learning for LLMs

Ten questions on the policy gradient theorem, PPO, reward modelling, RLHF, KL penalties, and the RL concepts that underpin modern LLM alignment and fine-tuning.