RLHF, DPO, and the reward models that turn a base modelโs raw capability into a model people actually want to use.