arXiv:2601.11578cs.CLcs.AI2026-01

用多智能体系统挖掘论文深层局限,提升科学严谨性。

Multi-Agent LLMs for Generating Research Limitations

  • 设计多角色智能体协同分析显性、隐性与同行评审视角的局限
  • 相较零样本模型,限制覆盖度提升15.51%(GPT-4o mini)
  • 适合需要深度评估研究缺陷的研究者或审稿人使用

识别和表述研究局限是科学透明与严谨的关键。然而,零样本大语言模型常生成浅显或泛化的局限陈述(如数据偏差或可推广性),通常重复作者已提过的局限,未能深入分析方法论漏洞或上下文空白。这一问题因作者常仅披露部分或琐碎的局限而加剧。本文提出一种多智能体大语言模型框架,用于生成实质性局限。该框架整合了OpenReview评论与作者自述局限,提供更强的基准;同时利用引用与被引文献捕捉更广泛的上下文薄弱点。不同智能体承担特定角色:部分提取显式局限,部分分析方法论缺口,部分模拟同行评审视角,另设引用智能体将研究置于更大文献体系中。一名裁判智能体优化输出,主控智能体整合为清晰集合。该结构可系统识别显性、隐性、评审导向及文献关联的局限。传统NLP指标(如BLEU、ROUGE、余弦相似度)依赖词元或嵌入重叠,常忽略语义相似性。为此,我们引入基于LLM作为裁判的逐项评估协议,更准确衡量覆盖度。实验表明,所提模型显著提升性能:RAG + 多智能体GPT-4o mini配置相比零样本基线提升15.51%覆盖度,而Llama 3 8B多智能体方案实现4.41%改进。

原文摘要 · Abstract (English)

Identifying and articulating limitations is essential for transparent and rigorous scientific research. However, zero-shot large language models (LLMs) approach often produce superficial or general limitation statements (e.g., dataset bias or generalizability). They usually repeat limitations reported by authors without looking at deeper methodological issues and contextual gaps. This problem is made worse because many authors disclose only partial or trivial limitations. We propose, a multi-agent LLM framework for generating substantive limitations. It integrates OpenReview comments and author-stated limitations to provide stronger ground truth. It also uses cited and citing papers to capture broader contextual weaknesses. In this setup, different agents have specific roles as sequential role: some extract explicit limitations, others analyze methodological gaps, some simulate the viewpoint of a peer reviewer, and a citation agent places the work within the larger body of literature. A Judge agent refines their outputs, and a Master agent consolidates them into a clear set. This structure allows for systematic identification of explicit, implicit, peer review-focused, and literature-informed limitations. Moreover, traditional NLP metrics like BLEU, ROUGE, and cosine similarity rely heavily on n-gram or embedding overlap. They often overlook semantically similar limitations. To address this, we introduce a pointwise evaluation protocol that uses an LLM-as-a-Judge to measure coverage more accurately. Experiments show that our proposed model substantially improve performance. The RAG + multi-agent GPT-4o mini configuration achieves a +15.51\% coverage gain over zero-shot baselines, while the Llama 3 8B multi-agent setup yields a +4.41\% improvement.

多智能体研究局限大模型评估科学严谨性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。