端到端实现说话人归属与时间戳转录,支持长达90分钟会议。
MOSS Transcribe Diarize Technical Report
- 统一多模态大模型,端到端完成说话人识别与时间标注。
- 128k上下文窗口,支持90分钟长音频输入,性能超越商用系统。
- 适用于会议转录、语音分析等需精确时间定位的场景。
说话人归属时间戳转录(SATS)旨在准确记录每个人说了什么以及发言的时间点,对会议转录尤为关键。现有SATS系统很少采用端到端架构,且受限于有限的上下文窗口、弱长时说话人记忆能力以及无法输出时间戳。为解决这些问题,我们提出MOSS Transcribe Diarize,一个统一的多模态大语言模型,以端到端方式联合执行说话人归属与时间戳转录。该模型在大量真实环境数据上训练,并具备128k上下文窗口,可处理长达90分钟的输入,展现出良好的可扩展性与强泛化能力。在多个公开及内部基准测试中,其表现均优于当前最先进的商用系统。
原文摘要 · Abstract (English)
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak long-range speaker memory, and the inability to output timestamps. To address these limitations, we present MOSS Transcribe Diarize, a unified multimodal large language model that jointly performs Speaker-Attributed, Time-Stamped Transcription in an end-to-end paradigm. Trained on extensive real wild data and equipped with a 128k context window for up to 90-minute inputs, MOSS Transcribe Diarize scales well and generalizes robustly. Across comprehensive evaluations, it outperforms state-of-the-art commercial systems on multiple public and in-house benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。