构建多模态表达数据集,提升音频视觉分割的理解与推理能力
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- 设计8类融合文本、语音、声音与视觉的多模态表达
- 在2104段视频上包含61095条需深层理解的指代表达
- 适合研究多模态推理与跨模态理解的学者参考
指代式音视频分割(RAVS)近年取得显著进展,但仍面临多模态信息融合与深层内容理解的挑战。为拓展该领域边界并推动后续研究,我们提出全新数据集OmniAVS,包含2,104段视频和61,095条多模态指代表达。其三大创新包括:(1)八种灵活组合文本、语音、声音与视觉线索的多模态表达;(2)强调对音频内容的深层理解,而不仅是存在检测;(3)在表达中融入复杂推理与世界知识。同时,我们提出OISA模型,利用多模态大语言模型(MLLM)实现复杂线索理解与基于推理的分割。大量实验表明,OISA在OmniAVS上优于现有方法,并在其他相关任务中表现竞争力。
原文摘要 · Abstract (English)
Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of RAVS and facilitate future research in this field, we propose Omnimodal Referring Audio-Visual Segmentation (OmniAVS), a new dataset containing 2,104 videos and 61,095 multimodal referring expressions. OmniAVS stands out with three key innovations: (1) 8 types of multimodal expressions that flexibly combine text, speech, sound, and visual cues; (2) an emphasis on understanding audio content beyond just detecting their presence; and (3) the inclusion of complex reasoning and world knowledge in expressions. Furthermore, we introduce Omnimodal Instructed Segmentation Assistant (OISA), to address the challenges of multimodal reasoning and fine-grained understanding of audiovisual content in OmniAVS. OISA uses MLLM to comprehend complex cues and perform reasoning-based segmentation. Extensive experiments show that OISA outperforms existing methods on OmniAVS and achieves competitive results on other related tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。