用多智能体辩论找出论文没说的局限,提升评审系统性。
Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique
- 设计五类专家角色,分领域深度分析论文漏洞。
- 在414篇论文上提升精度79%、覆盖率达11%以上。
- 适合审稿人或研究者做科学论文批判性评估。
随着科研文献激增,论文中未报告的局限性问题日益严重。本文提出 Tree-of-Concerns,一种基于多智能体的框架,通过五类专业化怀疑者角色,以类别特定视角并行展开结构化辩论,挖掘科学论文中未明言的局限。每个角色基于证据进行论证,再由评审委员会从五个角度重新评估存活论点,纠正类别偏差与严重性误判。在 ToC-Bench 数据集(含414篇论文、1,905个未陈述局限,来自审稿意见与引用批评)上的实验表明,该方法相较最强基线,精度提升79%,覆盖率提高11%,能有效揭示具体、有证据支持的关切点,辅助审稿人进行系统性评估。
原文摘要 · Abstract (English)
As scientific literature grows and papers increasingly under-report limitations, multi-agent LLMs offer a promising approach to systematically uncover these hidden failure modes. Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific papers. Each persona conducts structured, evidence-grounded argumentation, while a Panel Review mechanism re-evaluates each surviving claim from all five perspectives to correct category drift and severity miscalibration. Through experiments on ToC-Bench, our benchmark of 414 research papers with 1,905 unstated limitations, sourced from reviewer-reported weaknesses and follow-up citation critiques, we demonstrate that ToC improves precision by 79% and coverage by 11% relative to strongest baselines, surfacing specific, evidence-grounded concerns that support reviewers in systematic evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。