arXiv:2607.01527cs.SDcs.LG2026-07中稿 · INTERSPEECH 2026

提出可量化房间特征不确定性,仅用一次语音就能判断结果可信度。

Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score

论文配图:Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score
图 1 · 摘自论文原文
  • 构建基于声学脉冲响应的鲁棒嵌入空间,对抗语音内容干扰
  • 通过分布离散度校准不确定性得分,波形与频谱扰动下表现一致
  • 无需下游任务监督,单次语音即可实现可信预测筛选

从混响语音中提取的房间嵌入常不可靠:语音内容和录音退化会改变表征,即使说话人、房间及源-接收器几何结构不变,也会降低下游任务性能。我们提出一种框架,学习对语音内容变化具有鲁棒性的房间嵌入,并从混响语音中无监督地生成表示级不确定性得分。嵌入锚定于结构化的房间脉冲响应(RIR)潜在空间,利用基于Kullback-Leibler(KL)对齐的多视角数据结构进行训练;多正例对比项进一步提升鲁棒性。设计轻量级不确定性头,基于退化嵌入的分布离散度进行校准,并通过排序目标优化。在波形级与频谱级扰动下,该得分与表示离散度一致,可在仅需单次语音输入的情况下实现有效选择性预测。

原文摘要 · Abstract (English)

Room embeddings derived from reverberant speech are often unreliable: speech content and recording degradation can alter the representation even when speaker, room, and source-receiver geometry remain unchanged, degrading downstream task performance. We propose a framework that learns room embeddings robust to speech-content variation and a representation-level uncertainty score from reverberant speech without downstream-task supervision. The embedding is anchored to a structured room impulse response (RIR) latent space and trained using a multi-view data structure with Kullback-Leibler (KL)-based alignment; a multi-positive contrastive term further refines robustness. A lightweight uncertainty head is calibrated using the dispersion of corruption-induced embeddings and optimized with a rank-based objective. Across waveform- and spectrogram-level corruptions, the score is consistent with representation dispersion and enables effective selective prediction while requiring only a single utterance at inference.

语音处理不确定性估计嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。