arXiv:2601.21463cs.SDcs.AI2026-01被引 5

用音频大模型统一检测语音编辑并定位篡改内容,提升真实场景适应性。

Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs

  • 将语音编辑检测转为文本生成任务,联合推理编辑类型与位置。
  • 在140小时多样的真实编辑数据集上,检测准确率显著优于现有方法。
  • 引入概率提示和声学一致性损失,增强模型对真实音频证据的感知能力。

现有语音编辑检测(SED)数据集主要依赖人工拼接或有限编辑操作,多样性不足且难以覆盖真实场景。当前方法依赖帧级监督来识别可观察的声学异常,难以处理内容被完全删除的编辑。为此,我们提出一个基于音频大语言模型(Audio LLMs)的统一框架,将语音编辑检测与内容定位结合。首先构建AiEdit数据集(约140小时),涵盖使用先进端到端语音编辑系统生成的添加、删除、修改操作,提供更贴近现实的评测基准。在此基础上,将SED重构为结构化文本生成任务,实现编辑类型识别与内容定位的联合推理。为增强生成模型对声学证据的锚定,提出先验增强提示策略,注入由帧级检测器提取的词级概率线索。同时引入声学一致性感知损失,显式约束潜在空间中正常与异常声学表示的分离。实验表明,该方法在检测与定位任务上均持续优于现有方法。

原文摘要 · Abstract (English)

Existing speech editing detection (SED) datasets are predominantly constructed using manual splicing or limited editing operations, resulting in restricted diversity and poor coverage of realistic editing scenarios. Meanwhile, current SED methods rely heavily on frame-level supervision to detect observable acoustic anomalies, which fundamentally limits their ability to handle deletion-type edits, where the manipulated content is entirely absent from the signal. To address these challenges, we present a unified framework that bridges speech editing detection and content localization through a generative formulation based on Audio Large Language Models (Audio LLMs). We first introduce AiEdit, https://huggingface.co/datasets/JunXueTech/AiEdit, a large-scale bilingual dataset (approximately 140 hours) that covers addition, deletion, and modification operations using state-of-the-art end-to-end speech editing systems, providing a more realistic benchmark for modern threats. Building upon this, we reformulate SED as a structured text generation task, enabling joint reasoning over edit type identification, and content localization. To enhance the grounding of generative models in acoustic evidence, we propose a prior-enhanced prompting strategy that injects word-level probabilistic cues derived from a frame-level detector. Furthermore, we introduce an acoustic consistency-aware loss that explicitly enforces the separation between normal and anomalous acoustic representations in the latent space. Experimental results demonstrate that the proposed approach consistently outperforms existing methods across both detection and localization tasks.

语音编辑音频LLM内容定位生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。