Back to the index
05/ 55

SAFETY & SOCIETY

Alignment.

The work of making AI systems behave in ways that match human intentions, values, and constraints.

In plain words

Alignment is the research and engineering effort to make an AI system pursue appropriate objectives and behave as intended, including in situations its designers did not anticipate.

A closer look

A training objective is a measurable target, while a human intention is often complicated and unstated. A system rewarded for pleasing a user might learn to agree with a false claim. A system rewarded for completing a task might take an unacceptable shortcut. Alignment asks how to narrow these gaps.

Approaches include better training examples, feedback from people or other models, evaluations, and oversight. Researchers also study whether good behavior generalizes to unfamiliar settings and whether a system is optimizing for something different from its stated goal. Agreement about whose values matter is itself part of the problem.

In practice

AN EXAMPLE

A helpful assistant should correct an arithmetic mistake in your proposal even if praising it would earn a better immediate rating. Helpfulness includes honesty and respect for relevant constraints.

A useful distinction

Alignment is not simply obedience or a list of forbidden answers. Following every instruction can conflict with other people’s interests, truthfulness, or the purpose of the task. It is also not a solved property certified by one test.

Sources & further reading

Ji et al. — AI Alignment: A Comprehensive Survey (opens in a new tab)