arXiv:2605.12682cs.AI2026-05

让AI从少量对话中学会用户隐性偏好,提升决策人性化水平

Learning Transferable Latent User Preferences for Human-Aligned Decision Making

论文配图:Learning Transferable Latent User Preferences for Human-Aligned Decision Making
图 1 · 摘自论文原文
  • 从极少对话中学习可迁移的自然语言偏好规则
  • 在多个任务和环境下显著提升决策对齐度,推理成本更低
  • 适合需要个性化、低交互智能决策的应用场景

大语言模型(LLMs)被广泛用于各类应用中的推理模块,但常难以生成符合人类偏好的解决方案。人类对齐决策需同时考虑明确目标与影响模糊情境处理的隐性偏好。现有方法或依赖大量重复用户交互,或无法跨任务泛化隐性偏好,实用性受限。本文提出CLIPR(对话式偏好与推理学习框架),使LLM在有限交互下推断用户隐性偏好,并生成可操作、可迁移的自然语言规则。这些规则通过自适应反馈迭代优化,应用于分布内与分布外的模糊任务,在三个数据集及用户研究中均显著优于现有方法,提升对齐度并降低推理开销。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as reasoning modules in many applications. While they are efficient in certain tasks, LLMs often struggle to produce human-aligned solutions. Human-aligned decision making requires accounting for both explicitly stated goals and latent user preferences that shape how ambiguous situations should be resolved. Existing approaches to incorporating such preferences either rely on extensive and repeated user interactions or fail to generalize latent preferences across tasks and contexts, limiting their practical applicability. We consider a setting in which an LLM is used for high-level reasoning and is responsible for inferring latent user preferences from limited interactions, which guides downstream decision making. We introduce CLIPR (Conversational Learning for Inferring Preferences and Reasoning), a framework that learns actionable, transferable natural language rules that represent latent user preferences from minimal conversational input. These rules are iteratively refined through adaptive feedback and applied to both in-distribution and out-of-distribution ambiguous tasks across multiple environments. Evaluations on three datasets and a user study show that CLIPR consistently outperforms existing methods in improving alignment and reducing inference costs.

人机对齐偏好学习对话系统可迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。