arXiv:2606.26842eess.AScs.HC2026-06

开源工具voxmap-studio让语音分说话人标注更高效,还能量化标注成本。

voxmap-studio: An open-source speaker diarization annotation tool with built-in cost instrumentation

论文配图:voxmap-studio: An open-source speaker diarization annotation tool with built-in cost instrumentation
图 1 · 摘自论文原文
  • 用自动分说话人结果做初始假设,减少手动绘制发言段
  • 记录每步编辑操作和耗时,可定量比较不同辅助方式效果
  • 适合语音标注研究者、数据清洗团队使用

标注说话人分离数据成本高昂,但现有工具很少度量这一成本。我们提出voxmap-studio,一个基于React的开源标注工具,集成pyannote分说话人生态。其画布由快速步长加速的分说话人引擎初始化,使标注员只需修正预测结果而非从零开始绘制发言段。工具记录编辑操作次数和时间作为第一类输出,支持量化比较不同辅助形式的实际帮助效果。导出需逐段人工确认,并通过注入的“幽灵”注意力检测防止未经验证的自动结果被当作真实标签发布。在9个AMI音频文件的初步研究中,完全手动标注成本最高且准确率最低;自动初始化将工作重心从创建发言段转向修正;突出不确定片段的策略在小样本中成本最低。工具及成本度量功能均已开源。

原文摘要 · Abstract (English)

Labeling speaker diarization data is costly, yet annotation tools rarely measure that cost. We present voxmap-studio, an open-source, React-based diarization annotation tool integrated with the pyannote-based diarization ecosystem. Its canvas is initialized by a fast stride-accelerated diarization engine so that the annotator corrects a hypothesis rather than drawing every speaker turn by hand, and the tool records annotation cost - typed edit-operation counts and time - as a first-class output, enabling quantitative comparison of how much different forms of assistance actually help. Export is gated on per-segment human confirmation and guarded by injected "phantom" attention checks, which prevent unverified automatic output from being released as ground truth. In a preliminary study on nine AMI audio files, unassisted manual annotation was the costliest and least accurate, and automatic initialization shifted the work from creating turns to correcting them; highlighting uncertain segments gave the lowest cost in our small sample. The tool and its instrumentation are open source.

语音标注工具开发成本度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。