针对对话中口语化谣言,构建多轮音频数据集并提升事实核查效果。
Context-Aware Multimodal Claim Verification in Spoken Dialogues

- 融合上下文音频与对话文本的多模态模型,捕捉发言者间互动逻辑。
- 仅用前序对话上下文即可接近离线验证性能,适合实时审核场景。
- 音频信息在文本模型受干扰时贡献最大,凸显语音语调关键作用。
每天数百万用户从播客和直播中接收未经核查的言论,而这些口头信息常通过对话积累可信度,其真伪不仅取决于事实,更取决于话语如何被框架、强化或未被反驳。然而现有事实核查研究多聚焦孤立文本,忽视对话音频价值。本文提出MAD2——一个包含1000组双人对话、3368个可验证陈述及约10小时音频的新基准数据集,并设计一种校准的多模态融合方法:结合上下文感知音频编码器与对话感知文本模型。实验表明,在不同场景下引入对话上下文均能提升验证效果,但增益程度依赖任务类型;仅使用前序上下文即能达到接近离线性能,支持实时内容审核;当文本模型因额外上下文出现波动时,音频模态贡献最大。整体而言,对话结构比误导性表述框架对验证影响更大。
原文摘要 · Abstract (English)
Every day, millions absorb claims from podcasts and streams that no fact-checker ever sees. Spoken misinformation is built through conversation, where credibility comes not from facts alone but from how claims are framed, reinforced, or left unchallenged across turns. Yet fact-checking has focused on isolated text, leaving dialogue audio under-studied. We introduce MAD2, a new Multi-turn Audio Dialogues benchmark for spoken claim verification, containing 1,000 two-speaker dialogues with 3,368 check-worthy claims and approximately 10 hours of audio, and propose calibrated multimodal fusion of a context-aware audio encoder and a dialogue-aware text model. Across settings, adding dialogue context improves verification, but the gains depend on scenario type. Using only preceding context often matches offline performance, supporting live-moderation settings, and audio contributes most when transcript-based models are destabilized by additional context. Overall, conversational structure matters more for verification than misinformation framing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。