arXiv:2603.05459cs.CLcs.DB2026-03被引 1

构建首个口语化个体辩论语料库,支持多任务研究。

DEBISS: a Corpus of Individual, Semi-structured and Spoken Debates

  • 收集真实口语辩论数据,具半结构化特征。
  • 覆盖语音转写、说话人分离等多类标注。
  • 适合研究辩论分析与交互式对话系统者使用。

辩论在学习、工作、日常讨论、电视政治辩论及社交媒体中无处不在,应用广泛。由于辩论形式多样、结构复杂,现有语料库数量稀少且难以覆盖其多样性。为此,本文提出 DEBISS 语料库:一个包含口语化、个体化、半结构化辩论的语料集,涵盖语音转写、说话人分离、论点挖掘及辩手质量评估等多任务标注,为自然语言处理中的辩论研究提供基础数据支持。

原文摘要 · Abstract (English)

The process of debating is essential in our daily lives, whether in studying, work activities, simple everyday discussions, political debates on TV, or online discussions on social networks. The range of uses for debates is broad. Due to the diverse applications, structures, and formats of debates, developing corpora that account for these variations can be challenging, and the scarcity of debate corpora in the state of the art is notable. For this reason, the current research proposes the DEBISS corpus: a collection of spoken and individual debates with semi-structured features. With a broad range of NLP task annotations, such as speech-to-text, speaker diarization, argument mining, and debater quality assessment.

语料库辩论分析语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。