根据查询动态选择音视频模态,提升广播视频中人物检索精度
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

- 通过跨模态得分一致性判断当前查询中哪些模态有效
- 在BBC Rewind数据集上达到94.2%的P@1,优于单一模态和固定融合
- 适合处理真实广播场景中音/视频缺失不全的问题
在真实广播档案中进行语音与人脸联合人物检索时,目标可能仅可听不可见、可见但不可听,或两者皆有。若强行融合缺失模态的得分,反而会引入噪声,导致性能低于最优单模态系统。本文提出一种查询自适应框架,通过跨模态得分一致性检测活跃模态:当双模态均有效时,一模态检索结果在另一模态上也应得分高;模态缺失时该一致性消失。基于此设计的分类器实现89%检测准确率。在包含超过12,000段视频的BBC Rewind语料库上,该系统取得94.2%的P@1,优于仅说话人(82.9%)、仅人脸(93.4%)及固定融合(90.0%),恢复了64%与拥有真实模态标签的理论最优系统(96.6%)之间的差距。
原文摘要 · Abstract (English)
When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may be heard but unseen, seen but unheard, or both. Fusing scores from an absent modality injects noise, degrading precision below the best unimodal system. We propose a query-adaptive framework that detects active modalities via cross-modal score consistency: when both modalities are active, files retrieved by one also score highly on the other; this agreement breaks down when a modality is absent. Classifiers driven by these cross-modal features achieve 89% detection accuracy. On the BBC Rewind corpus (with over 12,000 broadcast videos) the adaptive system attains 94.2% P@1, outperforming speaker-only (82.9%), face-only (93.4%), and fixed fusion (90.0%), recovering 64% of the gap to an oracle with ground-truth modality labels (96.6%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。