arXiv:2608.02171cs.AI2026-08

评测大模型代理如何根据用户历史行为隐式理解偏好并执行任务

From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents

论文配图:From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
图 1 · 摘自论文原文
  • 基于长期交互记录构建动态偏好评测基准,捕捉用户行为中的隐含线索
  • 现有顶尖模型在复杂场景下仍难以实现有效个性化,表现普遍不佳
  • 提出新框架通过全局检索与轨迹对齐提升多任务一致性,适合个性化应用开发者

大型语言模型已推动自主代理能力不断提升,但个性化仍是其实用化的关键。现有评测多依赖静态偏好快照、固定交互日志或预定义用户画像的问答任务,无法反映用户偏好的动态演变,也忽视了基于偏好的任务执行——我们称之为知识到行动的差距。为此,我们提出IBA-Bench,一个基于包含噪声、隐含信号和时间不一致性的纵向交互历史构建的隐式行为对齐评测基准。不同于以往工作,IBA-Bench评估代理能否从历史交互中推断出隐式约束并完成任务。我们进一步提出IBA-Agent框架,通过广泛检索与轨迹级对齐解决优先级冲突。在IBA-Bench上的实验表明,当前顶尖大模型代理在个性化方面仍面临显著挑战,而IBA-Agent在九个应用领域中显著提升了复杂场景下的行为对齐效果。

原文摘要 · Abstract (English)

Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have begun evaluating personalization in agents, but they largely rely on static preference snapshots, fixed interaction logs, or question answering over predefined user profiles. Such designs fail to capture the complexity of evolving user preferences and neglect preference-conditioned task execution-a discrepancy we term as the knowledge-to-action gap. To address this challenge, we introduce IBA-Bench, a benchmark for implicit behavioral alignment constructed from longitudinal interaction histories that contain noise, implicit cues, and temporal inconsistencies. Unlike prior work, IBA-Bench evaluates whether an agent can execute tasks while satisfying implicit user constraints inferred from historical interactions. We further propose IBA-Agent, an agent framework that reconciles conflicting priorities through broad retrieval and trajectory-level alignment. Experiment results on IBA-Bench show that effective personalization remains a significant challenge for state-of-the-art LLM agents, and the proposed IBA-Agent substantially improves behavioral alignment in complex scenarios across nine application domains.

个性化代理行为对齐评测基准隐式偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。