arXiv:2607.21635cs.LG2026-07中稿 · KDD被引 2

提出评估个性化大模型代理的新标准,关注用户状态随时间变化下的行为一致性。

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

  • 设计四条件评估框架:时间干预、持久状态、跨维度影响、用户状态差异
  • 审计现有基准发现无一满足全部条件,暴露评测空白
  • 建议最小化基准设计与报告指标,适用于个性化代理开发与测试

个性化代理会随用户演化记忆、技能、工具配置和策略状态。现有基准多孤立评估各项能力:工具调用测试固定API下的表现,记忆测试召回或遗忘,安全测试静态合规性。本文认为个性化代理评估需新范式:在不同持久用户状态下重复相同时间干预,并测量故障如何在组件间传播。我们提出四条核心条件:明确的时间干预、干预期间的持久状态、引发跨维度效应、用户状态差异。通过显式筛选标准对公开基准进行定向审计,发现虽有若干接近案例,但在严格操作化定义下,未找到满足全部四条件的协议。本研究以有限文献覆盖为基础,定位此为聚焦性评测缺口。论文提出最小化基准设计及候选报告指标,为未来个性化代理评估提供具体设计要求,并以指标作为实现该要求的报告工具。

原文摘要 · Abstract (English)

Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or forgetting, and safety benchmarks test static policy compliance. We argue that personal-agent evaluation requires a different protocol: replaying the same temporal intervention across different persistent user-conditioned states and measuring how failures propagate across agent components. We formalize this requirement as four conditions: explicit temporal intervention, persistent state across the intervention, induced cross-dimensional effects, and variation in user-conditioned state. A focused audit of public benchmark protocols selected by explicit inclusion criteria identifies several close cases. Under our explicitly narrow operationalization, we did not find a protocol in that audited set satisfying all four conditions. This claim is scoped as a focused gap analysis with bounded literature coverage. This position paper proposes a minimal benchmark design and candidate reporting metrics for user-conditioned adaptation. The result is a concrete design requirement for future personal-agent evaluation, with metrics used as reporting tools for that requirement.

个性化代理评估基准时间干预用户状态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。