arXiv:2605.10936cs.CV2026-05被引 3

让大模型学会理解用户独有的视觉信息,实现真正个性化助手。

Personal Visual Context Learning in Large Multimodal Models

论文配图:Personal Visual Context Learning in Large Multimodal Models
图 1 · 摘自论文原文
  • 构建用户视觉记忆库,动态整合个人化视觉证据。
  • 在多任务测试中显著优于传统提示方法,提升个性化推理能力。
  • 适合研究个性化AI助手、可穿戴设备智能交互的学者与开发者。

随着智能眼镜等可穿戴设备将大型多模态模型(LMMs)融入用户的持续第一视角视觉流,这些模型演变为真正个人助理的关键在于视觉个性化:即对佩戴者独有的视觉信息进行推理。我们将其定义为个人视觉上下文学习(Personal VCL),即在提示阶段利用用户特定视觉上下文解决个性化问题的能力。为系统评估该能力,我们提出了 Personal-VCL-Bench,一个全面捕捉个体、物体和行为层面个人视觉世界的基准。对前沿LMMs的分析揭示了显著的上下文利用率差距,表明利用视觉证据及聚合多个视觉观测的机制仍严重缺乏研究。基于此,我们提出智能体上下文银行(Agentic Context Bank),一种强大的推理时基线,将用户视觉上下文结构化为自我优化的记忆库,并采用查询自适应证据选择。该基线在多个任务和模型骨干上持续优于标准上下文提示范式,为未来个性化LMMs提供了可行路径。

原文摘要 · Abstract (English)

As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models into true personal assistants hinges on visual personalization: the ability to reason over visual information unique to the wearer. We formalize this capability as Personal Visual Context Learning (Personal VCL), the prompt-time capability of using user-specific visual context to resolve personalized queries. To systematically evaluate this, we present Personal-VCL-Bench, a comprehensive benchmark capturing the personal visual world across persons, objects, and behaviors. Our analysis of frontier LMMs identifies a profound context utilization gap, revealing that the mechanisms for leveraging visual evidence, as well as aggregating multiple visual observations, remain critically understudied. Motivated by these findings, we propose the Agentic Context Bank, a strong inference-time baseline that structures a user's visual context into a self-refining memory bank and employs query-adaptive evidence selection. Our baseline approach consistently improves over standard context prompting regimes across tasks and evaluated backbones, demonstrating a practical path towards future personalized LMMs.

多模态个性化视觉上下文可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。