In chapter 2 we taught the model by "imitating the answer key" one token at a time. But the qualities that make an AI assistant genuinely usable —
correct, polite, not making things up, not sliding into English — have no answer key to imitate, and cannot be written directly as a loss function.
This chapter is the most classical answer to that problem: RLHF (Reinforcement Learning from Human Feedback) with PPO.
We will train a real reward model from Thai preference pairs, then write the PPO loop ourselves from scratch in about 120 lines,
and close with my favorite experiment in the series: taking off the KL leash and watching the model cheat the reward live.
This is deliberately the heaviest chapter of the series, because chapter 4 (DPO) and chapter 5 (GRPO)
both start from this chapter's equations and each choose a different piece to delete.