arXiv:2608.04095cs.AIcs.CL2026-08

评测大模型在金融场景中长期维护用户偏好记忆的能力

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

论文配图:FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
图 1 · 摘自论文原文
  • 基于金融事件构建真实用户行为轨迹,强制模型追踪偏好变化
  • 7个主流大模型在长时记忆任务上准确率不足47%,多选题仅39%
  • 发现摘要记忆易丢失偏好信号,简单检索反而更有效

大型语言模型代理在金融顾问等高风险领域日益应用,但其能否长期维持并更新个性化用户模型仍不明确。现有个性化记忆基准多聚焦事实记忆或依赖弱约束的模型生成轨迹,对事件驱动的偏好适应研究不足。本文提出FinPerMA,一个基于真实金融事件的基准,评估模型在冻结的长期投资者轨迹中的个性化记忆能力。其生成流程结合确定性、理论驱动的影响规则、受控的LLM叙述和自动化质量筛选;通过'冲击后检查点'判断模型是否将重大事件纳入持久用户模型。在2,994个问题、276个角色的测试中,七个前沿大模型及最多七种记忆配置均未达饱和:无完整上下文配置的综合准确率不超过0.47,多选题准确率约39%。归因分析显示,基于摘要的记忆常保留事实细节却丢失偏好信号,导致简单检索反而优于专用记忆系统,且冲击后差距扩大。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.

个性化记忆大模型评估金融智能长期记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。