arXiv:2607.24958eess.AScs.SD2026-07

针对执法记录仪音频难题,构建可落地的对话智能框架。

Towards Operational Conversational Intelligence: A Speech Intelligence Framework

论文配图:Towards Operational Conversational Intelligence: A Speech Intelligence Framework
图 1 · 摘自论文原文
  • 分路处理:语音分离与转写并行,再融合结果
  • 在复杂录音下实现高精度说话人区分与文字标注
  • 模块化设计适合后续扩展,对执法场景特别实用

穿戴式摄像头(BWC)音频面临环境噪声大、录制条件不一、多人重叠讲话等挑战,导致自动转录和说话人标注困难。本文提出一种双路径对话智能框架:先对原始BWC音频进行预处理,将处理流程分为说话人归因分支和语音识别分支,最后融合输出。说话人归因分支使用降噪前端(DeepFilterNet)、语音活动检测(VAD)以及NVIDIA的多尺度说话人归因解码器(MSDD)与TitaNet嵌入;转写分支采用响度归一化,结合WhisperX(Large-v3)进行强制对齐和概率引导的语音分割。最终通过词级时间重叠匹配,将每个识别出的词分配给对应说话人段。在从公开的美英警方执法记录中构建的定制化数据集上评估,实验表明任务特定声学调制和概率引导语音分割显著提升在复杂条件下的说话人归因、转录及词级说话人标注性能。所提模块化架构为未来说话人感知对话智能系统提供可扩展基础。

原文摘要 · Abstract (English)

Body-worn camera (BWC) audio presents unique challenges including high ambient noise, variable recording conditions, and multiple overlapping speakers that make automated transcription and speaker labeling challenging. We propose a dual-path conversational intelligence framework that preprocesses raw BWC audio, separates the processing pipeline into a diarization branch and an ASR branch, and fuses their outputs. The diarization branch uses a denoising front-end (DeepFilterNet), voice activity detection (VAD), and NVIDIA's Multi-Scale Speaker Diarization Decoder (MSDD) with TitaNet embeddings. The transcription branch uses loudness normalization and WhisperX (Large-v3) with forced alignment and probability-guided speech segmentation. Finally, word-level speaker attribution is performed by assigning each recognized word to the speaker segment with the greatest temporal overlap. We evaluate the proposed framework on a curated body-worn camera dataset constructed from publicly available U.S. and U.K. police body-worn camera recordings. Experimental results demonstrate that task-specific acoustic conditioning and probability-guided speech segmentation improve speaker diarization, transcription, and word-level speaker attribution under challenging body-worn camera recording conditions. The proposed modular architecture provides an extensible foundation for future speaker-aware conversational intelligence systems.

对话智能说话人归因语音识别执法记录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。