RL & Alignment
Reinforcement Learning for LLMs
Ten questions on the policy gradient theorem, PPO, reward modelling, RLHF, KL penalties, and the RL concepts that underpin modern LLM alignment and fine-tuning.
- 10 questions · single best answer
- Explanation revealed after each answer
- Score shown at the end — no data stored