arXiv:2607.02504cs.CLcs.AI2026-07中稿 · ICML

用大模型提升长剧对白角色识别准确率

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

论文配图:Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas
图 1 · 摘自论文原文
  • 基于大推理模型构建多模态证据聚合框架
  • 在532K对白数据上实现比基线显著提升
  • 特别擅长处理声纹不可靠的短对白

长篇电视剧对视频理解构成重大挑战,其中理清复杂剧情往往依赖于说话人识别——即准确将每个语音片段归因到对应角色。本文提出两项主要贡献:(1) 构建了DramaSR-532K大规模基准数据集,包含超过900个角色的532,000条标注对白,需融合听觉、语言和视觉线索进行识别;(2) 提出DramaSR-LRM,一种基于大推理模型(LRM)的鲁棒方法。该方法可自主调用多模态工具,聚合上下文证据,实现高保真角色归属。实验表明,DramaSR-LRM显著优于现有基线,尤其在声纹信息不可靠的短对白场景下表现突出。所有数据与代码将在项目页面公开:https://www.github.com/198808xc/DramaSR-LRM。

原文摘要 · Abstract (English)

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character. In this paper, we advance this field through two primary contributions. (1) We introduce \textbf{DramaSR-532K}, a large-scale benchmark comprising 532K annotated dialogue lines across more than 900 unique characters, necessitating the integration of auditory, linguistic, and visual cues for speaker recognition. (2) We propose \textbf{DramaSR-LRM}, a robust approach built upon a large reasoning model (LRM). DramaSR-LRM is designed to autonomously aggregate contextual evidence via multimodal tool-use, synthesizing diverse inputs to achieve high-fidelity attribution. Experimental results demonstrate that DramaSR-LRM significantly outperforms existing baselines, particularly on short utterances where acoustic biometrics are inherently unreliable. \textit{All the data and code will be made publicly available at the project page: https://www.github.com/198808xc/DramaSR-LRM.}

说话人识别多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。