提出动态人机协作对齐新范式,强调交互中涌现偏好而非静态模仿。
Align AI to Dynamic Human-AI Workflows

- 用轨迹级视角建模人机行为共演化,突破静态偏好假设。
- 揭示人机协作中不确定性加剧与协调难题的新挑战。
- 呼吁融合机器学习与社会科学,推动交互式对齐研究。
当前对齐方法多基于静态人类偏好表示来模拟人类行为,难以捕捉真实人机交互中动态、上下文依赖的特性。本文主张从静态模仿转向交互互补的对齐范式,认为偏好应在互动中自然涌现,对齐不应仅满足既有偏好。我们通过轨迹级视角揭示现有对齐方法的不足,强调人机行为随时间共同演化的本质。由于现有机器学习框架未充分建模此类动态,本文整合跨学科研讨会洞见,借鉴人类协作的社会科学理论,指出人机系统放大了协作复杂性,引入新的不对称性,使不确定性推理更难,并带来全新协调挑战。最后,我们提出一项研究议程,旨在发展在持续交互中实现对齐的AI系统,需融合机器学习与社会及决策科学。
原文摘要 · Abstract (English)
Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions. In this paper, we argue for a shift from static and emulative to interactive and complementary alignment, where preferences emerge through interaction and alignment is defined not by satisfying preferences alone. We first formalize this gap by contrasting existing alignment with a trajectory-level view in which human and model behavior co-evolve over time. Because these interaction dynamics have not been adequately captured within existing ML formulations, we ground this perspective in insights from an interdisciplinary workshop. We draw on lessons from social-science accounts of human-human collaboration and then argue that human-AI systems amplify these dynamics, introducing new asymmetries that make reasoning about uncertainty harder and introduce new coordination challenges. Based on these lessons and new challenges, we conclude by outlining a research agenda for developing AI systems that align with humans in interaction, requiring an interdisciplinary synthesis of machine learning and the social and decision sciences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。