A separate model trained to score how good or bad a response is, used to guide another model's training toward producing higher-scoring outputs.
Reward models are central to RLHF — instead of a human rating every single output during training, a reward model learns to predict what a human would rate highly.
One of 60 free AI glossary terms
Plain-language definitions for the AI jargon you'll actually run into — no email needed, ever, for this section.
Browse the Full Glossary →