arXiv:2510.17132cs.LGcs.AI2025-10被引 4

测试大模型能否通过对话发现用户隐藏偏好,揭示其个性化能力边界。

Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction

  • 设计三阶段对话任务评估模型挖掘隐性偏好的能力
  • 不同任务下准确率从32%到98%波动,受主题和属性数量影响大
  • 适合研究个性化对话、智能助手与用户建模的开发者

大型语言模型在生成通用文本方面表现优异,但在需要用户特定偏好的场景(如推荐餐厅或规划旅行)中,其通用性成为瓶颈。用户通常不会明确表达所有偏好,许多关键信息处于隐性状态,需通过对话推断。本文提出首个系统性基准,评估模型在多轮对话中发现并利用隐藏用户属性的能力。基准涵盖三个渐进式真实场景:经典20个问题游戏、个性化问答和个性化摘要生成。所有任务均采用三代理框架(用户、助手、裁判),支持逐轮评估信息获取与适应能力。结果显示,尽管模型能通过对话揭示隐性信息,但成功率随任务复杂度、话题和隐藏属性数量显著变化,范围为32%至98%。该基准为研究个性化交互中的隐性信息发现提供了统一框架,表明有效偏好推断仍是构建真正自适应AI系统的关键挑战。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel at producing broadly relevant text, but this generality becomes a limitation when user-specific preferences are required, such as recommending restaurants or planning travel. In these scenarios, users rarely articulate every preference explicitly; instead, much of what they care about remains latent, waiting to be inferred. This raises a fundamental question: Can LLMs uncover and reason about such latent information through conversation? We address this problem by introducing a unified benchmark for evaluating latent information discovery - the ability of LLMs to reveal and utilize hidden user attributes through multi-turn interaction. The benchmark spans three progressively realistic settings: the classic 20 Questions game, Personalized Question Answering, and Personalized Text Summarization. All tasks share a tri-agent framework (User, Assistant, Judge) enabling turn-level evaluation of elicitation and adaptation. Our results reveal that while LLMs can indeed surface latent information through dialogue, their success varies dramatically with context: from 32% to 98%, depending on task complexity, topic, and number of hidden attributes. This benchmark provides the first systematic framework for studying latent information discovery in personalized interaction, highlighting that effective preference inference remains an open frontier for building truly adaptive AI systems.

个性化对话偏好推断多轮交互评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。