用大模型生成模糊指纹,提升纯文本对话中说话人识别准确率
Speaker Fuzzy Fingerprints: Benchmarking Text-Based Identification in Multiparty Dialogues
- 引入说话人专属标记与上下文感知建模,利用大模型提取模糊指纹
- 在Friends和大爆炸理论数据集上准确率分别达70.6%和67.7%
- 可替代全微调且更易解释,适合研究对话中身份混淆问题
仅基于文本的说话人识别面临挑战,因传统语音特征不可用。现有方法多依赖传统手段,效果有限。本文探索利用大型预训练模型生成的模糊指纹,结合说话人专属标记与上下文感知建模,显著提升识别精度,在Friends数据集上达到70.6%,Big Bang Theory数据集上为67.7%。结果表明,模糊指纹可在较少隐藏单元下逼近全微调性能,且更具可解释性。此外,我们分析了模糊语句并提出检测无主语对话行的机制。研究揭示了关键挑战,并为未来改进提供洞见。
原文摘要 · Abstract (English)
Speaker identification using voice recordings leverages unique acoustic features, but this approach fails when only textual data is available. Few approaches have attempted to tackle the problem of identifying speakers solely from text, and the existing ones have primarily relied on traditional methods. In this work, we explore the use of fuzzy fingerprints from large pre-trained models to improve text-based speaker identification. We integrate speaker-specific tokens and context-aware modeling, demonstrating that conversational context significantly boosts accuracy, reaching 70.6% on the Friends dataset and 67.7% on the Big Bang Theory dataset. Additionally, we show that fuzzy fingerprints can approximate full fine-tuning performance with fewer hidden units, offering improved interpretability. Finally, we analyze ambiguous utterances and propose a mechanism to detect speaker-agnostic lines. Our findings highlight key challenges and provide insights for future improvements in text-based speaker identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。