arXiv:2605.26072cs.LG2026-05

通过智能生成查询提升偏好学习效率,减少标注成本。

Active Query Synthesis for Preference Learning

论文配图:Active Query Synthesis for Preference Learning
图 1 · 摘自论文原文
  • 在连续空间中最大化互信息,自动合成最优查询
  • 新模型识别模糊比较,提高反馈可靠性
  • 适用于文本摘要与机器人控制等复杂场景

高效学习用户偏好对现代决策系统至关重要,但通常依赖昂贵的标注数据。主动学习可降低这一成本,但传统方法因基于池的评估计算开销大。多数方法假设所有查询反馈同样可靠,却忽视了相似或完全不同的项目间的对比会产生模糊、低置信度响应。为此,我们提出一种新的置信度感知响应模型,显式处理这些模糊比较。为克服基于池的评估计算瓶颈,我们提出主动查询合成框架 Info-Synth,通过在连续空间中最大化基于互信息的目标来生成最优查询。此外,我们设计两种策略——Pair M-dist 和 Pair Opt-dist,使 Info-Synth 在有限查询池下仍能有效选择查询。我们在合成偏好学习、约束性文本摘要数据集以及模拟移动机器人控制器增益的主观连续空间调优任务中验证了该框架的通用性与性能。

原文摘要 · Abstract (English)

Efficient learning of user preferences is crucial for many modern decision making systems but typically requires costly labeled data. Active learning reduces this cost, yet standard methods are computationally expensive due to pool-based evaluation. Further, most methods assume all query feedback is equally reliable, ignoring that pairwise queries between nearly identical or entirely dissimilar items yield ambiguous, low-confidence responses. To address the issue of feedback reliability, we introduce a novel confidence aware response model that explicitly accounts for these ambiguous comparisons. To overcome the computational bottleneck of pool-based evaluation, we propose an active query synthesis framework, Info-Synth that generates optimal queries by maximizing a mutual information-based objective within a continuous space. Moreover, we propose two strategies, Pair M-dist and Pair Opt-dist, that extend Info-Synth to select effective queries even when restricted to finite query pools. We demonstrate our framework's versatility and performance across synthetic preference learning, constrained text summary datasets, and subjective, continuous-space controller gain tuning for a simulated mobile robot.

偏好学习主动学习查询生成机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。