arXiv:2512.21616cs.CV2025-12KDD被引 6

提出无需训练的框架TAME,让多模态大模型在长对话中持续记忆并理解个性化对象变化。

TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant

  • TAME用双记忆机制区分处理个性化概念的时序与持久变化
  • 在长对话任务中表现最优,支持动态更新的个性化交互
  • 适合需要长期记忆和个性化服务的智能助手场景

多模态大语言模型(MLLM)个性化是实现针对特定实体(即个性化概念)进行个性化对话的关键问题。然而,现有方法与基准主要关注简单的视觉识别与文本替换(如将"一只黄狗"替换为"你的狗Mochi"),忽略了长上下文对话能力。理想的个性化MLLM助手应能与人类进行长上下文对话,并通过过往对话历史持续提升体验质量。为此,我们提出了首个长上下文MLLM个性化评估基准LCMP,用于评估模型对个性化概念变化的感知能力及生成适配响应的能力。作为强基线,我们引入一种无需训练且具备状态感知的TAME框架。TAME通过双记忆机制差异化管理每个个性化概念的时间与持久变化。此外,TAME采用全新的无需训练的检索-对齐增强生成(RA2G)范式,通过一个对齐步骤从多记忆检索的知识中提取语境适配信息,以应对复杂真实用户查询。在LCMP上的实验表明,TAME表现最佳,在长上下文场景中展现出显著且不断演进的交互体验。

原文摘要 · Abstract (English)

Multimodal Large Language Model (MLLM) Personalization is a critical research problem that facilitates personalized dialogues with MLLMs targeting specific entities (known as personalized concepts). However, existing methods and benchmarks focus on the simple, context-agnostic visual identification and textual replacement of the personalized concept (e.g., "A yellow puppy" -> "Your puppy Mochi"), overlooking the ability to support long-context conversations. An ideal personalized MLLM assistant is capable of engaging in long-context dialogues with humans and continually improving its experience quality by learning from past dialogue histories. To bridge this gap, we propose LCMP, the first Long-Context MLLM Personalization evaluation benchmark. LCMP assesses the capability of MLLMs in perceiving variations of personalized concepts and generating contextually appropriate personalized responses that reflect these variations. As a strong baseline for LCMP, we introduce a novel training-free and state-aware framework TAME. TAME endows MLLMs with double memories to manage the temporal and persistent variations of each personalized concept in a differentiated manner. In addition, TAME incorporates a new training-free Retrieve-then-Align Augmented Generation (RA2G) paradigm. RA2G introduces an alignment step to extract the contextually fitted information from the multi-memory retrieved knowledge to the current questions, enabling better interactions for complex real-world user queries. Experiments on LCMP demonstrate that TAME achieves the best performance, showcasing remarkable and evolving interaction experiences in long-context scenarios.

多模态大模型个性化长对话无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。