arXiv:2606.06837eess.AScs.LG2026-06中稿 · Interspeech 2026

提出抗捷径的实时口语风格检测框架,提升面试场景可信度。

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

论文配图:SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails
图 1 · 摘自论文原文
  • 通过统一预处理与缝合感知采样避免数据偏差
  • 8秒窗口下外部测试集AUC达0.971,显著优于基准
  • 适合需可靠口语检测的面试系统部署

scripted与spontaneous语音检测在面试监管中具有潜力,但现有基准性能可能因语料身份、信道条件和录音伪影等捷径而被夸大,而非真实反映说话风格。我们提出SEAM,一种面向实时剧本化检测的抗捷径框架,结合统一预处理、缝合感知采样、非语音增强和紧凑的DistilHuBERT骨干网络。使用8秒窗口,在外部面试域评估集上模型达到0.971 ± 0.004的ROC-AUC。移除防捷径组件后,内部指标提升但外部性能急剧下降,表明存在捷径学习现象。训练后量化将模型体积压缩至41.8MB,外部性能损失极小。结果表明,鲁棒的实时剧本化检测不仅依赖骨干网络,更取决于抗捷径的数据设计与评估。代码与模型检查点已开源。

原文摘要 · Abstract (English)

Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus identity, channel conditions, and recording artifacts rather than speaking style itself. We present SEAM, a shortcut-aware framework for real-time scriptedness detection that combines uniform preprocessing, seam-aware sampling, non-speech augmentation, and a compact DistilHuBERT backbone. With 8s windows, the model achieves 0.971 +- 0.004 ROC-AUC on an external interview-domain evaluation set. Removing the shortcut-prevention components improves internal held-out metrics but sharply reduces external performance, indicating shortcut learning. Post-training quantization reduces the model footprint to 41.8MB with little loss in external performance. The results demonstrate that robust real-time scriptedness detection depends not only on the backbone, but on shortcut-aware data design and evaluation. We release code and model checkpoints.

语音分析面试系统抗捷径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。