用多模型辩论模拟学术讨论,精准识别科学论文引用。
TurQUaz at CheckThat! 2025: Debating Large Language Models for Scientific Web Discourse Detection
- 让多个大模型像学术会议一样辩论,由主持人协调达成共识。
- 在识别科学论文引用上表现最佳,准确率排名第一。
- 适合需要高可信度科学内容检测的研究者和平台。
本文介绍我们在 CheckThat! 2025 科学网络话语检测任务(任务4a)中的工作。提出一种新型议会辩论方法,通过多个大语言模型(LLMs)模拟结构化学术讨论,判断推文是否包含:(i) 科学主张,(ii) 对科学文献的引用,或 (iii) 科学实体提及。我们探索三种辩论方式:单次辩论(两模型对立争辩,一模型裁判)、团队辩论(每方多模型协作)和议会辩论(多专家模型共同讨论,由主持模型引导)。最终选择议会辩论作为主模型,因其在开发测试集上表现最优。尽管该方法在识别科学主张(10个方案中第8名)和科学实体提及(第9名)上未进入前列,但在检测科学文献引用方面排名第一。
原文摘要 · Abstract (English)
In this paper, we present our work developed for the scientific web discourse detection task (Task 4a) of CheckThat! 2025. We propose a novel council debate method that simulates structured academic discussions among multiple large language models (LLMs) to identify whether a given tweet contains (i) a scientific claim, (ii) a reference to a scientific study, or (iii) mentions of scientific entities. We explore three debating methods: i) single debate, where two LLMs argue for opposing positions while a third acts as a judge; ii) team debate, in which multiple models collaborate within each side of the debate; and iii) council debate, where multiple expert models deliberate together to reach a consensus, moderated by a chairperson model. We choose council debate as our primary model as it outperforms others in the development test set. Although our proposed method did not rank highly for identifying scientific claims (8th out of 10) or mentions of scientific entities (9th out of 10), it ranked first in detecting references to scientific studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。