arXiv:2608.20777cs.CL2026-08中稿 · EMNLP

用多智能体辩论找出论文没说的局限,提升评审系统性。

Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique

  • 设计五类专家角色,分领域深度分析论文漏洞。
  • 在414篇论文上提升精度79%、覆盖率达11%以上。
  • 适合审稿人或研究者做科学论文批判性评估。

随着科研文献激增,论文中未报告的局限性问题日益严重。本文提出 Tree-of-Concerns,一种基于多智能体的框架,通过五类专业化怀疑者角色,以类别特定视角并行展开结构化辩论,挖掘科学论文中未明言的局限。每个角色基于证据进行论证,再由评审委员会从五个角度重新评估存活论点,纠正类别偏差与严重性误判。在 ToC-Bench 数据集(含414篇论文、1,905个未陈述局限,来自审稿意见与引用批评)上的实验表明,该方法相较最强基线,精度提升79%,覆盖率提高11%,能有效揭示具体、有证据支持的关切点,辅助审稿人进行系统性评估。

原文摘要 · Abstract (English)

As scientific literature grows and papers increasingly under-report limitations, multi-agent LLMs offer a promising approach to systematically uncover these hidden failure modes. Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific papers. Each persona conducts structured, evidence-grounded argumentation, while a Panel Review mechanism re-evaluates each surviving claim from all five perspectives to correct category drift and severity miscalibration. Through experiments on ToC-Bench, our benchmark of 414 research papers with 1,905 unstated limitations, sourced from reviewer-reported weaknesses and follow-up citation critiques, we demonstrate that ToC improves precision by 79% and coverage by 11% relative to strongest baselines, surfacing specific, evidence-grounded concerns that support reviewers in systematic evaluation.

多智能体科学评审局限提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。