arXiv:2607.20485cs.AIcs.CL2026-07中稿 · ICML

让大模型真正理解用户真实期待,提升对话匹配度。

Expectation Alignment of Language Models for Real-World User Expectations

论文配图:Expectation Alignment of Language Models for Real-World User Expectations
图 1 · 摘自论文原文
  • 从真实用户交互中提取语义丰富的期望,构建新评测基准
  • 现有模型在满足和预判用户期望上表现不佳,存在显著偏差
  • 提出轻量级框架LENS,显式建模用户期待,提升响应契合度

大型语言模型在标准基准测试中表现优异,但其是否真正满足用户实际期待仍不明确。现有评估方法依赖模型启发式、专家评分或用户模拟,难以捕捉真实人类期待的多样性与细微差别,导致模型看似能力强劲,实则与用户所求错位。本文首次系统研究真实场景下用户对大模型的期待,提出一种提取语义丰富期望的规范流程,并构建基于真实用户期待的ExpectBench基准。分析显示,当前大模型难以满足甚至预见用户期望,揭示了人机对齐的根本性错位。基于此,我们提出LENS——一种轻量级隐式期望感知响应生成框架,使模型能内化用户期待,生成更契合的回应,持续提升期望满足度,凸显显式建模用户期待对实现真实人机对齐的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model heuristics, expert rubrics, or user simulation, fail to capture the diversity and subtlety of real human expectations, causing models to appear competent while misaligning with what users actually seek. We present the first systematic study of user expectations in real-world LLM interactions, proposing a principled procedure to extract semantically rich expectations and introducing ExpectBench, a benchmark grounded in real user expectations. Analyses reveal that current LLMs struggle to satisfy and anticipate what users hope to obtain, highlighting a fundamental source of misalignment. Building on these observations, we propose LENS, a lightweight latent expectation-aware response generation framework. LENS enables models to internalize user expectations and generate better-aligned responses, consistently improving expectation satisfaction and underscoring the importance of explicitly modeling user expectations for realistic human-AI alignment.

大模型对齐用户期望响应生成评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。