用人类反馈迭代优化语音提取,让模型自动修复错误段落。
Neural Speech Extraction with Human Feedback
- 用户标记错误段落,系统仅修复指定区域,保留其他部分。
- 基于噪声功率的掩码训练的模型效果最好,与人工标注一致。
- 22人测试显示,经反馈优化的输出更受用户青睐。
我们提出了首个利用人类反馈进行迭代优化的神经目标语音提取(TSE)系统。用户可标记TSE输出中的特定片段,生成编辑掩码,系统仅优化标记区域而保持未标记部分不变。由于难以收集大规模人工标注错误数据,我们采用多种自动化掩码函数生成合成数据集,并在每种数据上训练模型。评估表明,使用基于噪声功率(单位:dBFS)的掩码和概率阈值训练的模型表现最优,且与人工标注高度一致。一项包含22名参与者的实验显示,用户更偏好经过反馈优化的输出。结果表明,人机协同优化是提升神经语音提取性能的有前景方向。
原文摘要 · Abstract (English)
We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then improves the marked sections while preserving unmarked regions. Since large-scale datasets of human-marked errors are difficult to collect, we generate synthetic datasets using various automated masking functions and train models on each. Evaluations show that models trained with noise power-based masking (in dBFS) and probabilistic thresholding perform best, aligning with human annotations. In a study with 22 participants, users showed a preference for refined outputs over baseline TSE. Our findings demonstrate that human-in-the-loop refinement is a promising approach for improving the performance of neural speech extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。