arXiv:2606.28845cs.CV2026-06中稿 · ECCV

让AI更精准识别用户独特概念,避免干扰信息。

Personalizing MLLMs via Reinforced Multimodal Reference Game

论文配图:Personalizing MLLMs via Reinforced Multimodal Reference Game
图 1 · 摘自论文原文
  • 设计对比性强化对话机制,让模型自己优化描述
  • 在三个基准上达到当前最佳性能,泛化能力强
  • 适合需要个性化视觉理解的场景,如智能助手

个性化多模态大模型旨在从视觉数据中识别用户独有的概念并提供定制化响应。尽管先前工作表明概念描述和推理有益于该任务,但多模态大模型的描述常包含状态、上下文等无关信息,可能干扰目标概念在视觉相似项中的唯一识别。有效的个人概念描述应准确、有区分性且无冗余细节。为此,我们提出强化参考游戏(RRG),一种通过新颖的强化多模态参考游戏促进区分性描述的学习框架。多模态大模型在对比性博弈设定中同时扮演说话者与听者角色,目标是有效传达关于目标概念的区分性信息。我们的方法对硬正例(同一概念的不同视角)与硬负例(视觉相似但不同的概念)构建可验证的对比奖励。实证结果表明,RRG在三个个性化基准上的多个任务中达到当前最优表现,且能泛化至未见领域,优于基于概念描述和个人化专用强化学习框架的现有方法。代码与模型将发布于项目页面。

原文摘要 · Abstract (English)

Personalizing Multimodal Large Language Models (MLLMs) aims to recognize users' unique concepts from visual data and provide personalized responses. Although prior work has shown the benefit of concept descriptions and reasoning for this task, MLLM descriptions often include information, such as state and context, that does not help and may in fact hinder the unique identification of the target concept among other visually similar items. Effective descriptions of personal concepts should instead be accurate, discriminative, and free of distracting details. To achieve such descriptions, we introduce Reinforced Reference Game (RRG), a learning framework that promotes discriminative descriptions through a novel reinforced multimodal reference game. The MLLM plays both the roles of speaker and listener in a contrastive game setting, whose goal is to effectively communicate discriminative information about a target concept. Our approach formulates a verifiable contrastive reward over hard positives (dissimilar views of the same concept) and hard negatives (visually similar but different concepts). Empirically, RRG achieves state-of-the-art across multiple tasks on three personalization benchmarks. RRG generalizes to unseen domains and outperforms existing methods based on concept descriptions and personalization-specific RL frameworks. We will release code and models in the project page.

多模态模型个性化强化学习视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。