arXiv:2510.10077cs.CL2025-10被引 1

让AI更懂用户潜在意图,提升个性化偏好对齐效果。

A-IPO: Adaptive Intent-driven Preference Optimization

  • 通过意图模块挖掘用户输入背后的深层需求,融入奖励函数优化模型响应。
  • 在真实与对抗场景下,胜率最高提升24.8%,意图一致性提升54.6%。
  • 适合追求个性化、鲁棒性对话系统的开发者和研究者。

人类偏好具有多样性和动态性,受地域、文化和社会因素影响。现有对齐方法如直接偏好优化(DPO)常默认多数意见,忽视少数观点,难以捕捉提示中的潜在意图。为此,我们提出自适应意图驱动偏好优化(A-IPO)。A-IPO引入意图模块,推断每个用户提示背后的潜在意图,并将其显式纳入奖励函数,增强优选模型回应与用户深层意图的一致性。理论上和实证上均证明,加入意图-响应相似项可使偏好边际提升(对数几率正向偏移λΔsim),显著区分优选与非优选回应。我们构建了两个新基准:Real-pref与Attack-pref,以及扩展版的GlobalOpinionQA-Ext,用于评估现实与对抗场景下的偏好对齐性能。实验表明,通过显式建模多元意图,A-IPO实现包容性偏好优化的同时,显著提升对抗鲁棒性。全面评估显示,A-IPO持续优于现有基线,在关键指标上取得显著提升:Real-pref上最高+24.8胜率与+45.6意图一致性;Attack-pref上最高+38.6响应相似度与+52.2防御成功率;GlobalOpinionQA-Ext上最高+54.6意图一致性得分。

原文摘要 · Abstract (English)

Human preferences are diverse and dynamic, shaped by regional, cultural, and social factors. Existing alignment methods like Direct Preference Optimization (DPO) and its variants often default to majority views, overlooking minority opinions and failing to capture latent user intentions in prompts. To address these limitations, we introduce \underline{\textbf{A}}daptive \textbf{\underline{I}}ntent-driven \textbf{\underline{P}}reference \textbf{\underline{O}}ptimization (\textbf{A-IPO}). Specifically,A-IPO introduces an intention module that infers the latent intent behind each user prompt and explicitly incorporates this inferred intent into the reward function, encouraging stronger alignment between the preferred model's responses and the user's underlying intentions. We demonstrate, both theoretically and empirically, that incorporating an intention--response similarity term increases the preference margin (by a positive shift of $λ\,Δ\mathrm{sim}$ in the log-odds), resulting in clearer separation between preferred and dispreferred responses compared to DPO. For evaluation, we introduce two new benchmarks, Real-pref, Attack-pref along with an extended version of an existing dataset, GlobalOpinionQA-Ext, to assess real-world and adversarial preference alignment. Through explicit modeling of diverse user intents,A-IPO facilitates pluralistic preference optimization while simultaneously enhancing adversarial robustness in preference alignment. Comprehensive empirical evaluation demonstrates that A-IPO consistently surpasses existing baselines, yielding substantial improvements across key metrics: up to +24.8 win-rate and +45.6 Response-Intention Consistency on Real-pref; up to +38.6 Response Similarity and +52.2 Defense Success Rate on Attack-pref; and up to +54.6 Intention Consistency Score on GlobalOpinionQA-Ext.

偏好优化意图建模鲁棒对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。