arXiv:2509.23251cs.MMcs.SD2025-09被引 3

多智能体系统提升视频音频对齐与关键片段检索效率。

XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System

  • 四阶段协作框架:感知、规划、执行、反思,优化多模态理解流程。
  • 在2685个问答上测试,现有多模态模型在质量感知和对齐上表现不佳。
  • 首个涵盖AI生成内容的音视频理解评测集,适合评估真实与生成场景能力。

本文提出XGC-AVis,一个增强多模态大模型(MLLMs)音视频时序对齐能力并提升关键视频片段检索效率的多智能体框架,包含感知、规划、执行、反思四个阶段。我们进一步构建了XGC-AVQuiz,首个全面评估MLLMs在真实世界与AI生成场景下理解能力的基准测试,包含2,685个问答对,覆盖20项任务。其两大创新在于:1)AIGC场景扩展:包含2,232个视频,涵盖1,102个专业生成内容(PGC)、753个用户生成内容(UGC)和377个AI生成内容(AIGC),覆盖10大领域与53个细粒度类别;2)质量感知维度:在识别、定位、推理之外,新增需融合低层感官与高层语义的能力,评估音视频质量、同步性与连贯性。实验表明当前MLLMs在质量感知与时序对齐任务中表现不足。XGC-AVis无需额外训练即可提升性能,在两个基准上得到验证。

原文摘要 · Abstract (English)

In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the efficiency of retrieving key video segments through 4 stages: perception, planning, execution, and reflection. We further introduce XGC-AVQuiz, the first benchmark aimed at comprehensively assessing MLLMs' understanding capabilities in both real-world and AI-generated scenarios. XGC-AVQuiz consists of 2,685 question-answer pairs across 20 tasks, with two key innovations: 1) AIGC Scenario Expansion: The benchmark includes 2,232 videos, comprising 1,102 professionally generated content (PGC), 753 user-generated content (UGC), and 377 AI-generated content (AIGC). These videos cover 10 major domains and 53 fine-grained categories. 2) Quality Perception Dimension: Beyond conventional tasks such as recognition, localization, and reasoning, we introduce a novel quality perception dimension. This requires MLLMs to integrate low-level sensory capabilities with high-level semantic understanding to assess audio-visual quality, synchronization, and coherence. Experimental results on XGC-AVQuiz demonstrate that current MLLMs struggle with quality perception and temporal alignment tasks. XGC-AVis improves these capabilities without requiring additional training, as validated on two benchmarks.

多模态音视频对齐AI生成内容评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。