arXiv:2607.24821cs.MMcs.CV2026-07

评测音视频联合编辑能力,发现主流模型难以兼顾跨模态一致性和内容保真。

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

论文配图:AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
图 1 · 摘自论文原文
  • 构建145个真实视频的多模态编辑基准,覆盖196条联合指令与2688项细粒度检查点。
  • 现有模型在执行跨模态指令时,保真度和现实感普遍不足,尤其影响非目标区域。
  • 提出AVE-Agent框架,通过任务分解与自我反思提升音视频协同编辑效果。

尽管基于指令的视频编辑发展迅速,但真实视频中音频与视觉信号高度耦合,修改一模态通常需同步调整另一模态。现有基准主要评估静音片段的视觉变换或孤立音频编辑,对复杂音视频联合编辑及跨模态一致性关注不足。我们提出AVE-Compass,一个包含145个精选源视频、196条音视频耦合编辑指令和2688项细粒度检查项的综合评测基准。该基准通过检查清单式MLLM评分与专用现实感评分标准,评估指令遵循性、保真度维持、现实感及编辑意图实现,并辅以自动化跨模态、视频与音频指标。大量实验表明,当前先进模型在执行跨模态指令并保持非目标内容方面仍表现不佳。为此,我们进一步提出AVE-Agent,一种模块化智能体框架,可将复杂指令分解为依赖子任务,通过自省与评估反馈迭代优化编辑结果。AVE-Agent显著提升了指令执行率、保真度维持与音视频对齐能力,同时保持良好感知质量。

原文摘要 · Abstract (English)

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.

音视频编辑多模态评估跨模态对齐智能体框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。