首个跨模态个人化基准,揭示模型在音频与视觉上的接地缺陷。
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

- 构建文本、图像、音频统一的个人化任务图谱,涵盖18项细粒度任务
- 提出校准准确率指标,同时评估正确响应与合理回避,覆盖无身份查询场景
- 发现大模型仍会幻觉,且标注成本限制微调扩展性,指导未来训练设计
尽管多模态大语言模型在文本、图像和音频方面取得进展,个人化研究仍以视觉-语言为主,缺乏涵盖文本、图像和音频的统一基准,且未系统考虑缺失身份场景或接地行为分析。本文提出 Omni-Persona,首个全面的跨模态个人化基准。将任务形式化为在人格模态图上的跨模态路由,包含4个任务组和18项细粒度任务,共约750个样本。为严谨诊断接地行为,提出校准准确率(Cal),联合评估正确接地与适当回避,并在统一框架内纳入无身份查询。实验发现:(i) 近期开源模型存在一致的音频-视觉接地差距,通过密集规则监督的强化学习视觉回复(RLVR)部分缓解;(ii) 召回率和参数规模并非充分诊断指标,强召回可能伴随无身份幻觉,更大模型未必提升Cal值,表明校准是独立评估维度;(iii) SFT受真实标注可扩展性限制。在当前奖励设计下,RLVR提升召回但未改善校准回避,使Cal值维持或低于基础模型。Omni-Persona 成为诊断跨模态个人化陷阱的框架,指导未来后训练与奖励设计。
原文摘要 · Abstract (English)
While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, with unified omnimodal benchmarking that jointly covers text, image, and audio still limited, and lacking the methodological rigor to account for absent-persona scenarios or systematic grounding studies. We introduce Omni-Persona, the first comprehensive benchmark for omnimodal personalization. We formalize the task as cross-modal routing over the Persona Modality Graph, encompassing 4 task groups and 18 fine-grained tasks across ~750 items. To rigorously diagnose grounding behavior, we propose Calibrated Accuracy (Cal), which jointly evaluates correct grounding and appropriate abstention, incorporating absent-persona queries within a unified evaluation framework. On our dedicated experiments, three diagnostic findings emerge: (i) recent open-weight models show a consistent audio-vs-visual grounding gap that RLVR partially narrows via dense rule-based supervision; (ii) recall score and parameter scale are incomplete diagnostics, since strong recall can coexist with absent-persona hallucination and larger models do not always achieve higher Cal, exposing calibration as a separate evaluation axis; and (iii) SFT is limited by the scalability of ground-truth annotation. RLVR improves recall but not calibrated abstention under our reward design, leaving Cal at or below the base model. Omni-Persona thus serves as a diagnostic framework that surfaces the pitfalls of omnimodal personalization, guiding future post-training and reward design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。