arXiv:2609.02242cs.AIcs.HC2026-09

让AI助手在用户评估能力有限时,主动选择能学习偏好的好提案。

Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

  • 提出可评估性感知的协作框架,用提案探测用户偏好与评估限制。
  • 在图任务中,新方法比不考虑评估成本的基线提升23%以上。
  • 适合需要高效学习用户偏好的人机协作场景,如智能设计辅助。

AI助手常通过提出候选方案供用户评估来协作。现有方法多关注提案质量或用户目标推断,假设用户能可靠评估任意提案,但现实中受限于认知能力,这一假设常失效。本文研究可评估性感知的提案规划,将提案既作为任务干预,也作为探测用户隐含偏好与评估约束的探针,利用信念更新指导后续提案。我们形式化该问题为ProSE——一个隐藏参数的序贯协助问题,并采用KL正则化的有界理性二元响应模型,其中接受决策权衡价值收益与距离相关的可评估性惩罚。分析表明,高接受率的提案与信息丰富的探针未必一致,仅追求接受率会系统性表现更差。我们实现的ProSE-Plan是一种深度为2的贝叶斯自适应规划器,根据可能的响应和响应引发的后验信念评分提案。在受控图模拟中,当评估成本为瓶颈时,ProSE-Plan显著优于无评估意识和短视基线;探针-确认消融实验证实其选择的信息丰富提案是简单方法所忽略的。结果表明,用户可评估性是人工智能协助中一个关键且互补的规划维度,不同于生成质量和偏好推断。

原文摘要 · Abstract (English)

AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can fail in practice because of bounded rationality. We study evaluability-aware proposal planning, where proposals serve both as task interventions and as probes for learning latent preferences and evaluation constraints, where the resulting belief updates then guide later proposals. We formalise this setting as ProSE, a hidden-parameter sequential assistance problem, and instantiate it with a KL-regularised bounded-rational binary response model in which acceptance trades off value gain against a distance-dependent evaluability penalty. Analysing the planning consequence of this likelihood reveals that likely accepted proposals and informative probes need not coincide, which explains why planners that only pursue acceptance systematically underperform. We operationalise ProSE with \textsc{ProSE-Plan}, a depth-2 Bayes-adaptive planner that scores proposals by possible responses and response-induced posterior beliefs. In controlled graph simulations, \textsc{ProSE-Plan} improves over evaluability-unaware and myopic baselines when evaluation cost is the bottleneck, and a probe-commit ablation confirms that our approach selects informative proposals that simpler methods miss. Our results thus identify user evaluability as a planning-relevant dimension of AI assistance, complementary to generation quality and preference inference.

人机协作偏好学习可评估性规划算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。