针对环境音与语音分别篡改的新型欺骗检测难题,提出三阶段级联框架。
EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge

- 三阶段级联设计:先判混合真伪,再用双分支特征融合检测篡改类型。
- 在CompSpoofV2数据集上达到0.8266的宏平均F1,排名挑战赛第二。
- 适合研究音频安全、对抗复杂伪造场景的工程师和研究人员。
真实场景中的语音欺骗已从仅语音伪造演变为更复杂的组件级设置,其中语音与环境声音可独立被操纵。为应对这一挑战,本文提出针对ESDD2挑战赛的环境感知三阶段级联框架EnvTriCascade。首先,混合一致性检测器提供二元先验,用于区分原始录音与篡改混合信号,以校准最终决策。其次,两个互补的五分类检测器分别利用SSLAM+XLS-R与EAT-large+XLS-R的表征,通过跨分支注意力门控分类器整合多分支特征。为增强对多样混合条件的鲁棒性,引入RawBoost数据增强。系统仅在官方CompSpoofV2数据集上训练,测试集上取得0.8266的宏平均F1分数,显著优于官方基线,位列挑战赛第二名。
原文摘要 · Abstract (English)
ADD in real-world scenarios has evolved from speech-only spoofing to more challenging component-level settings, where speech and environmental sounds may be independently manipulated. To tackle this, we propose EnvTriCascade, an Environment-Aware Tri-Stage Cascaded framework for the ESDD2 Challenge. First, a mix-consistency detector provides a binary prior to distinguish original recordings from manipulated mixtures, which calibrates the final decisions. Next, two complementary five-class detectors, leveraging SSLAM+XLS-R and EAT-large+XLS-R representations, extract robust multi-branch features integrated via a cross-branch attention-gated classifier. To enhance robustness against diverse mixing conditions, we incorporate RawBoost augmentation. Trained exclusively on the official CompSpoofV2 dataset, our system achieves a Macro-F1 score of 0.8266 on the test set, significantly outperforming the official baseline and ranking second in the challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。