Glossary

RLHF

Reinforcement learning from human feedback — training a model using human preference judgments to align its outputs.

← All terms