arXiv:2603.20422cs.CVcs.AI2026-03被引 8

提出流式视频个性化理解新任务,实现实时互动的AI助手能力

PEARL: Personalized Streaming Video Understanding Model

论文配图:PEARL: Personalized Streaming Video Understanding Model
图 1 · 摘自论文原文
  • 设计流式视频理解新任务与评估基准PEARL-Bench
  • 构建132段视频、2173个带时间戳的精细标注数据集
  • 提出无需训练的PEARL策略,适配多种模型提升实时个性化响应

人类对新概念的认知本质上是持续流式过程:我们不断识别新对象或身份,并随时间更新记忆。然而,当前多模态个性化方法主要局限于静态图像或离线视频,无法将连续视觉输入与即时现实反馈结合,难以实现未来AI助手所需的实时交互式个性化响应。为此,我们首次提出并正式定义了个性化流式视频理解(PSVU)这一新任务。为推动该方向研究,我们引入首个专为此场景设计的综合性基准PEARL-Bench,评估模型在精确时间戳下对个性化概念的响应能力,涵盖两种模式:(1) 帧级,聚焦离散帧中的特定人物或物体;(2) 新型视频级,关注跨连续帧展开的个性化动作。PEARL-Bench包含132段独特视频和2,173个细粒度标注,时间戳精确,通过自动化生成与人工验证相结合的流程保障概念多样性和标注质量。为应对这一挑战性新场景,我们进一步提出PEARL——一种即插即用、无需训练的策略,作为强大基线。在8种离线与在线模型上的广泛评估表明,PEARL在三种不同架构上均取得显著且一致的性能提升,证明其高效且鲁棒。我们希望本工作推动视觉语言模型个性化发展,并激发更多关于流式个性化AI助手的研究。代码已开源:https://github.com/Yuanhong-Zheng/PEARL。

原文摘要 · Abstract (English)

Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimodal personalization methods are largely limited to static images or offline videos. This disconnects continuous visual input from instant real-world feedback, limiting their ability to provide the real-time, interactive personalized responses essential for future AI assistants. To bridge this gap, we first propose and formally define the novel task of Personalized Streaming Video Understanding (PSVU). To facilitate research in this new direction, we introduce PEARL-Bench, the first comprehensive benchmark designed specifically to evaluate this challenging setting. It evaluates a model's ability to respond to personalized concepts at exact timestamps under two modes: (1) Frame-level, focusing on a specific person or object in discrete frames, and (2) a novel Video-level, focusing on personalized actions unfolding across continuous frames. PEARL-Bench comprises 132 unique videos and 2,173 fine-grained annotations with precise timestamps. Concept diversity and annotation quality are strictly ensured through a combined pipeline of automated generation and human verification. To tackle this challenging new setting, we further propose PEARL, a plug-and-play, training-free strategy that serves as a strong baseline. Extensive evaluations across 8 offline and online models demonstrate that PEARL achieves state-of-the-art performance. Notably, it brings consistent PSVU improvements when applied to 3 distinct architectures, proving to be a highly effective and robust strategy. We hope this work advances vision-language model (VLM) personalization and inspires further research into streaming personalized AI assistants. Code is available at https://github.com/Yuanhong-Zheng/PEARL.

视频理解个性化流式处理AI助手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。