arXiv:2607.00368cs.CL2026-07

提出行为评估框架,验证大模型部署后记忆是否真实有效。

Beyond Perplexity: A Behavioral Evaluation Framework for Deployment-Memory Claims in LLM Test-Time Training

论文配图:Beyond Perplexity: A Behavioral Evaluation Framework for Deployment-Memory Claims in LLM Test-Time Training
图 1 · 摘自论文原文
  • 区分流适应、桥接内化与部署行为学习三类记忆,对应不同证据层级
  • 在稀疏事实场景中,LoRA微调降低损失但自由回忆为零,暴露代理指标缺陷
  • 适合关注模型真实记忆能力的研究者和评测人员

大语言模型测试时训练(TTT)常依赖局部代理指标:在近期标记、检索上下文、目标领域数据或可验证任务尝试上更新模型,再通过困惑度、未来标记损失、长上下文表现或奖励进行评估。这些指标适用于流适应、领域适应、上下文压缩及奖励驱动的改进,但对更常见且日益重要的部署后助手记忆、个性化或稀疏后部署学习等能力,证据力较弱。这类能力需行为证据,如后期回忆、改写鲁棒性、保留性、局部性、冲突处理以及原始支持上下文移除后的下游应用。本文提出一种行为评估框架,将TTT记忆声称与支持证据对齐。框架包含两个部分:一个声明校准的证据阶梯,区分流/领域适应、桥接内化与部署时行为学习;以及匹配显式记忆基线、互斥失败类别的评估协议。通过审计近期TTT与记忆相关工作,并在稀疏新事实设置下构建受控诊断实验,发现单步LoRA更新在三个Qwen3模型规模上均降低支持与答案损失,但生成的自由回忆始终为零,揭示了代理改进与部署行为之间的可测量差距。该框架为作者与评估者提供了明确标准,以确保记忆声称与实证证据一致。

原文摘要 · Abstract (English)

Large language model test-time training (TTT) is often evaluated through local proxy metrics: models are updated on recent tokens, retrieved context, target-domain data, or verifiable task attempts, and then judged by perplexity, future-token loss, long-context performance, or reward. These metrics are well matched to claims about stream adaptation, domain adaptation, context compression, and reward-backed test-time improvement. They are weaker evidence, however, for a capability that TTT results are increasingly used to motivate: deployed assistant memory, personalization, or sparse post-deployment learning, which instead requires behavioral evidence such as later recall, paraphrase robustness, retention, locality, conflict handling, and use in downstream actions after the original support context is removed. We introduce a behavioral evaluation framework that calibrates TTT memory claims to the evidence that supports them. It has two components: a claim-calibrated evidence ladder that separates stream/domain adaptation, bridge internalization, and deployment-time behavioral learning; and an evaluation protocol with matched explicit-memory baselines and mutually exclusive failure categories. We validate the framework by auditing recent TTT and memory-adjacent work and by instantiating it as a controlled diagnostic in which, in a sparse nonce-fact setting, one-step LoRA updates lower support and answer loss across three Qwen3 model scales while generated free-form recall stays at zero, exposing a measurable gap between proxy improvement and deployment behavior. The framework gives authors and evaluators a concrete standard for aligning TTT memory claims with the evidence actually reported.

大模型记忆评估行为验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。