通过分析人类任务解决轨迹中的三类错位,提升AI对人类思维的模仿能力。
Addressing and Visualizing Misalignments in Human Task-Solving Trajectories
- 提出统一框架定义三类思维错位:意图表达缺失、动作低效、意图错误。
- 设计启发式算法检测错位,实验证明对齐意图可显著提升模型表现。
- 适合关注人机思维对齐、强化学习中行为模仿的研究者。
理解人类任务解决轨迹中的错位对于提升模仿人类推理的AI模型至关重要。本研究将此类错位分为三类:(1) 缺乏表达意图的功能,(2) 动作序列效率低下,(3) 无法解决问题的错误意图。我们首先在统一框架中形式化并定义这三类错位,随后提出一种启发式算法,用于检测ARCTraj轨迹中的错位,并进行层次化与定量分析。此外,基于该形式化方法,我们提出一种意图估计方法,以推断用户行为与意图间的缺失对齐。通过轨迹对齐训练,实验表明,基于人类任务解决轨迹训练的AI模型在模仿人类推理方面表现更优。基于层次分析与实验,我们强调了轨迹-意图对齐的重要性,并验证了意图对齐训练的有效性。
原文摘要 · Abstract (English)
Understanding misalignments in human task-solving trajectories is crucial for enhancing AI models trained to closely mimic human reasoning. This study categorizes such misalignments into three types: (1) lack of functions to express intent, (2) inefficient action sequences, and (3) incorrect intentions that cannot solve the task. To address these issues, we first formalize and define these three misalignment types in a unified framework. We then propose a heuristic algorithm to detect misalignments in ARCTraj trajectories and analyze their impact hierarchically and quantitatively. We also present an intention estimation method based on our formalism that infers missing alignment between user actions and intentions. Through trajectory alignment, we experimentally demonstrate that AI models trained on human task-solving trajectories improve performance in mimicking human reasoning. Based on hierarchical analysis and experiments, we highlight the importance of trajectory-intention alignment and demonstrate the effectiveness of intention-aligned training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。