arXiv:2505.18071cs.CLcs.AI2025-05被引 6

让大模型从用户行为中推断个性化偏好,提升推理效率与准确性

Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals

  • 构建扩展推理链,系统化分析用户交互历史中的隐含偏好信号
  • 相比基线模型,跨域测试平均提升15.49%,支持实时增量推理
  • 适用于需要动态理解用户偏好的个性化应用,如智能助手、推荐系统

大型语言模型(LLMs)在数学和编程等复杂任务中表现出色,但其归纳推理——从不完整证据中提炼通用规则的能力——仍待深入探索。本文聚焦个性化偏好推断这一关键挑战,该任务要求模型从分散的行为信号中提炼一致的偏好模式。为此,提出AlignXplore模型,利用扩展推理链实现对用户交互历史中行为信号的系统性偏好推断。该方法支持高效流式推理:新信号出现时可直接基于已有偏好描述进行更新,无需重新处理全部历史数据,同时支持迭代优化。通过合成数据冷启动训练结合在线强化学习,实验表明,AlignXplore在域内与域外基准上平均优于基线模型15.49%,且对不同输入格式与下游模型具备良好泛化能力。进一步分析揭示了人类式归纳推理模式的涌现,并确立了偏好推断学习的最佳实践。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated significant success in complex reasoning tasks such as math and coding. In contrast to these tasks where deductive reasoning predominates, inductive reasoning-the ability to derive general rules from incomplete evidence, remains underexplored. This paper investigates extended inductive reasoning in LLMs through the lens of personalized preference inference, a critical challenge in LLM alignment where current approaches struggle to capture diverse user preferences. The task demands strong inductive reasoning capabilities as user preferences are typically embedded implicitly across various interaction forms, requiring models to synthesize consistent preference patterns from scattered signals. We propose AlignXplore, a model that leverages extended reasoning chains to enable systematic preference inference from behavioral signals in users' interaction histories. Such explicit preference articulation enables efficient streaming inference: when new behavioral signals emerge, the model can directly build upon previously inferred preference descriptions rather than reprocessing historical signals from scratch, while also supporting iterative refinement to the inferred preferences. We develop AlignXplore by combining cold-start training based on synthetic data with subsequent online reinforcement learning. Through extensive experiments, we demonstrate that AlignXplore achieves substantial improvements over the backbone model by an average of 15.49\% on in-domain and out-of-domain benchmarks, while maintaining strong generalization ability across different input formats and downstream models. Further analyses establish best practices for preference inference learning through systematic comparison of reward modeling strategies, while revealing the emergence of human-like inductive reasoning patterns during training.

个性化推理行为建模偏好推断大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。