arXiv:2605.00495cs.SDcs.CV2026-05中稿 · CVPR

让无声视频自动生成带时间标签的音效,提升音频合成实用性。

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

论文配图:MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video
图 1 · 摘自论文原文
  • 联合生成音频与音效事件标签,避免逐级误差累积。
  • 在Greatest Hits数据集上,音效起始点检测准确率提升至75.0%。
  • 适合需要精准音效标注的影视音效自动化生产场景。

近期多模态生成技术已实现从无声视频生成高质量音频。实际应用如音效制作不仅需要生成的音频,还需明确标注声音类型与时间信息。现有方法通常对生成音频进行后处理声学事件检测,但易产生误差累积。为此,我们提出MMAudio-LABEL(基于潜在空间的音效标注)框架,基于基础音频生成模型,联合从无声视频生成音频和帧对齐的音效预测。我们在Greatest Hits数据集上评估了音效起始点检测与17类材质分类任务。相比基线,本方法将音效起始点检测准确率从46.7%提升至75.0%,材质分类准确率从40.6%提升至61.0%。结果表明,联合学习音频生成与事件预测可实现更可解释、更实用的视频到音频合成。

原文摘要 · Abstract (English)

Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the generated audio but also explicit sound event labels detailing the type and timing of sounds. One straightforward approach involves applying a standard sound event detection to the generated audio. However, this post-hoc pipeline is inherently limited, as it is prone to error accumulation. To address this limitation, we propose MMAudio-LABEL (LAtent-Based Event Labeling), an event-aware audio generation framework built on a foundational audio generation model as its backbone that jointly generates audio and frame-aligned sound event predictions from silent videos. We evaluate our method on the Greatest Hits dataset for onset detection and 17-class material classification. Our approach improves onset-detection accuracy from 46.7% to 75.0% and material-classification accuracy from 40.6% to 61.0% over baselines. These results suggest that jointly learning audio generation and event prediction enables a more interpretable and practical video-to-audio synthesis.

音频生成音效标注多模态联合学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。