arXiv:2605.06987cs.LGcs.GT2026-05被引 1

用响应时间提升大模型对多样偏好的对齐效果

Response Time Enhances Alignment with Heterogeneous Preferences

论文配图:Response Time Enhances Alignment with Heterogeneous Preferences
图 1 · 摘自论文原文
  • 引入响应时间作为辅助信号,建模为漂移扩散模型
  • 在仅单次选择数据下仍能准确估计群体平均偏好
  • 无需用户标识,适合匿名数据收集场景

大语言模型对齐人类偏好通常依赖聚合反馈构建单一奖励模型,但该方法假设所有标注者具有相同偏好,忽略了现实中标注者的高度异质性与匿名性。仅使用二元选择数据会根本性扭曲学习策略,使真实群体平均偏好不可识别。本文证明,通过在偏好数据中加入简单的响应时间信号,可恢复群体平均偏好的可识别性。基于漂移扩散模型(DDM)建模每项决策,提出一种新型一致估计器,即使在每个匿名标注者仅提供一次选择的极端情况下,也能渐近收敛至真实平均偏好。在合成与真实数据集上,该方法持续优于传统基线,而后者因偏差上限陷入性能瓶颈。响应时间几乎零成本记录,无需用户追踪或身份识别,为未来数据采集流程带来新机遇,显著提升社会价值。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this standard approach assumes that all labelers share the same underlying preferences, ignoring the fact that real-world labelers are highly heterogeneous and usually anonymous. Consequently, relying solely on binary choice data fundamentally distorts the learned policy, making the true population-average preference unidentifiable. To overcome this critical limitation, we demonstrate that augmenting preference datasets with a simple, secondary signal -- the user's response time -- can restore the identifiability of the population's average preference. By modeling each decision as a Drift-Diffusion Model (DDM), we introduce a novel, consistent estimator of heterogeneous preferences that successfully corrects the distortions of standard choice-only labels. We prove that our estimator asymptotically converges to the true average preference even in extreme cases where each anonymous labeler contributes only a single choice. Empirically, across both synthetic and real-world datasets, our method consistently outperforms standard baselines that otherwise fail and plateau at a bias floor. Because response times are essentially free to record and require zero user tracking or identification, our results bring promises and open up new opportunities for future data-collection pipelines to improve the social benefit without requiring user-level identifiers or repeated elicitations.

偏好对齐响应时间异质性无标识数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。