Alignment
The problem of making AI systems reliably pursue what their operators (and society) actually intend, including training techniques like RLHF and DPO, and the broader research field studying it.
The problem of making AI systems reliably pursue what their operators (and society) actually intend, including training techniques like RLHF and DPO, and the broader research field studying it.