arXiv:2605.26969cs.CLcs.AI2026-05

用动作重建评估推理路径,让模型生成更真实的人类行为模拟。

Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling

论文配图:Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling
图 1 · 摘自论文原文
  • 通过重建动作来评分推理链条,确保其能自然推导出行为
  • 在四个领域中比基线方法高出54.7%胜率,最优达70.0%
  • 生成的推理可跨模型迁移,提升用户建模效果

用户建模旨在利用语言模型(LMs)从过往上下文-动作对(如对话轮次)中模仿个体行为,适用于行为科学、人机协作和市场研究等场景。现有方法通过合成推理轨迹增强语料,通常基于上下文与动作联合条件生成,但此类方法属于事后合理化:推理轨迹必然为动作辩护,却未必反映真实的因果决策过程。本文提出Recon,利用动作重建评估推理轨迹的预测能力:给定上下文与候选推理,重建模型预测动作,重建精度决定推理质量。在四个领域中,Recon相较标准事后合理化基线Backward Synthesis实现54.7%胜率;进一步发现,以Recon奖励训练推理合成模型,可使下游用户建模性能提升至最高70.0%胜率。此外,Recon合成的推理具备跨模型迁移能力,且能超越重建模型本身提升用户建模表现。研究证明,事后合理化不足以支撑有效推理合成,真正有用且可解释的推理应能从上下文中自然引出行为。

原文摘要 · Abstract (English)

User modeling aims to use language models (LMs) to mimic an individual's behavior from a corpus of past context-action pairs (e.g., conversation turns), enabling the simulation of users in settings like behavioral science, human-AI collaboration, and market research. Recent approaches augment these corpora with synthesized reasoning traces, typically generated by conditioning on both context and action. However, such conditioning constitutes post-hoc rationalization rather than reasoning: the trace is guaranteed to justify the action, but may not encode the underlying latent causal decision paths. We propose Recon, which uses action reconstruction to score reasoning traces by their predictive power: given a context and candidate reasoning, a reconstruction model predicts the action, and reconstruction fidelity determines reasoning quality. Across four domains, Recon achieves a 54.7% win rate over Backward Synthesis, a standard post-hoc rationalization baseline. Further, we find that training a reasoning synthesis model with rewards derived from Recon improves downstream user modeling performance, achieving a win rate of up to 70.0% over baselines. We further show that Recon-synthesized reasoning transfers across models, and improves user modeling beyond the reconstruction model. Our work demonstrates that post-hoc rationalization is insufficient for reasoning synthesis, and that useful and interpretable reasoning should naturally elicit the action from the context.

用户建模推理合成重建评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。