arXiv:2607.14189cs.CVcs.SD2026-07被引 2

构建首个多参考音视频生成综合评估基准,推动跨模态内容创作研究

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

论文配图:MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
图 1 · 摘自论文原文
  • 提出可扩展的资产组合流水线,生成350个精细标注样本
  • 设计四维评估体系,14项子指标量化质量、一致性与指令遵循度
  • 融合自动评分与大模型重判机制,实现可审计的全面评估

多参考到音视频(MR2AV)生成旨在根据多个参考源和文本指令生成连贯的音视频内容。现有基准主要聚焦于文本驱动生成、单参考主体保真或孤立的音视频对齐,未覆盖新兴的MR2AV场景。相比传统设定,MR2AV要求模型在生成同步视听内容时联合推理多个参考源,不仅需忠实保留每个参考,还需正确绑定并组合多个实体形成连贯的视听事件。为此,本文提出MultiRef-Compass,一个统一的MR2AV评估基准。该基准包含通过可扩展、可控的资产组合流程构建的350个精心策划样本,涵盖多视角主体保真、多实体绑定及人-物-场景组合。为提供可解释评估,定义了包含四项维度(基础质量、参考一致性、音视频一致性、指令遵循)的评估协议,采用14项子指标。MultiRef-Compass集成自动评估指标与增强重判的多模态大模型作为评判者框架,支持可扩展且可审计的感知保真度与参考条件组合评估。对八种代表性MR2AV系统的广泛实验表明,各维度均存在显著提升空间,凸显综合性基准的重要性,并确立MultiRef-Compass作为未来MR2AV研究的基础。

原文摘要 · Abstract (English)

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises $350$ carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.

音视频生成多模态评估参考一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。