arXiv:2502.01402cs.CL2025-02中稿 · TheWebConf 2025被引 4

打造实时标注播客的工具,助力多语言事实核查研究。

Annotation Tool and Dataset for Fact-Checking Podcasts

  • 支持播客播放时实时标注关键陈述与上下文错误。
  • 结合Whisper转录与众包标注,构建高质量多语言数据集。
  • 适合多语言事实核查、自然语言处理研究者使用。

播客是网络上流行的媒体形式,内容多样且多语种,常包含未经验证的陈述。事实核查播客极具挑战性,需完成转录、标注和陈述验证,并保留口语内容的上下文细节。我们开发的工具通过在播放过程中实时标注,使用户可边听边标记关键陈述、陈述片段及上下文错误。该方法融合OpenAI Whisper等先进转录模型与众包标注,构建高质量数据集,用于微调XLM-RoBERTa等多语言Transformer模型,以实现陈述检测与立场分类。我们还发布了标注后的播客文本及样例标注,并提供初步实验结果。

原文摘要 · Abstract (English)

Podcasts are a popular medium on the web, featuring diverse and multilingual content that often includes unverified claims. Fact-checking podcasts is a challenging task, requiring transcription, annotation, and claim verification, all while preserving the contextual details of spoken content. Our tool offers a novel approach to tackle these challenges by enabling real-time annotation of podcasts during playback. This unique capability allows users to listen to the podcast and annotate key elements, such as check-worthy claims, claim spans, and contextual errors, simultaneously. By integrating advanced transcription models like OpenAI's Whisper and leveraging crowdsourced annotations, we create high-quality datasets to fine-tune multilingual transformer models such as XLM-RoBERTa for tasks like claim detection and stance classification. Furthermore, we release the annotated podcast transcripts and sample annotations with preliminary experiments.

事实核查播客多语言标注工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。