level: research
most ai alignment work assumes human preferences are stable targets to learn and optimize. this clashes with evidence from psychology and behavioral economics showing preferences are layered, dynamic, and shaped by context. as ai systems become more persistent and personalized, they increasingly influence what people value over time. the paper argues that ignoring this co-construction risks misalignment, because the very act of interacting with ai can change what users want.
the authors propose constructive alignment, a paradigm that reframes alignment as a control problem over evolving preference trajectories. they model preferences as layered state variables that shift under system influence. using a control-theoretic approach, they formalize how ai actions and interaction design jointly steer these trajectories. the goal is not just to satisfy current preferences but to guide their development in beneficial directions, avoiding manipulation or unintended value drift.
the framework draws on constructivist social theory to emphasize that preferences are not pre-existing but emerge through interaction. it suggests design principles for ai systems that respect user autonomy while acknowledging the inevitable influence of technology on human values. this shifts the alignment challenge from static optimization to dynamic stewardship, requiring new safety measures and evaluation methods that account for long-term preference change.
why it matters: it highlights that ai systems can reshape user values over time, demanding alignment methods that prevent harmful preference shifts.