Alignment

In one line

The work of making sure AI systems do what people actually intend and act in line with human values, rather than causing harm or pursuing the wrong goal.

Telling a computer exactly what you want is harder than it sounds. A system can follow the letter of an instruction while missing its spirit. Alignment research tries to close that gap, so AI tools are helpful, honest and avoid causing harm.

A simple example: if a cleaning robot were rewarded only for "no visible mess", it might learn to hide mess under the rug. For chatbots, alignment work includes teaching them to refuse dangerous requests, admit uncertainty and not mislead users.

The term is used at two scales. Day to day, it is about making current tools behave well. Some researchers also use it for the longer-term challenge of keeping far more capable future systems, such as AGI, under meaningful human control. People disagree about how big that future risk is, and also about whose values an AI should be aligned with.

Related: AI agent, Bias (in AI)