arXiv:2605.30947cs.CL2026-05中稿 · EMNLP

让AI读懂人文研究,通过多智能体精准抓取和绑定原始文献证据。

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

论文配图:Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship
图 1 · 摘自论文原文
  • 设计多智能体框架,将人文研究基本操作转化为可协作的智能体角色。
  • 在古汉语与拉丁语研究中,召回更多引用的原始文献证据,质量评分更高。
  • 适合需要严谨文献支撑的人文研究者,提升论文可信度与深度。

基于大模型的研究代理在科学与工程领域发展迅速,因其研究依赖可执行实验、代码与量化信号。而人文学科则依赖对原始资料的阐释性、证据支撑型论证,其价值在于忠实引述、可验证出处与深度解读。现有研究代理采用通用规划、工具使用与反思机制,使这些学术操作隐含而不显。为弥补这一差距,我们提出SPIRE(Scholarly-Primitives-Inspired Research Engine),一个将人文研究基本实践——即「学术原语」——转化为多尺度细读底座上协作角色的多智能体框架,该底座包含段落、上下文图社区与跨文本语义聚类。在古典汉语文本与古希腊-罗马拉丁文研究的同行评审论文基准上,SPIRE在引用原始文献证据方面显著优于LLM、RAG与通用代理基线,并在四项学术质量维度上获得更高盲评人类与大模型评分。检索层级与智能体消融实验表明,其优势源于精准检索、结构化证据选取与论断-证据绑定。代码与基准数据已开源至https://github.com/YatingPan/SPIRE。

原文摘要 · Abstract (English)

LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantitative signals. Humanities scholarship, however, requires interpretive, evidence-grounded arguments over primary sources, whose value rests on faithful quotation, verifiable provenance, and deep interpretation. Existing research agents use general planning, tool use, and reflection, leaving these scholarly operations implicit. To address this gap, we introduce SPIRE (Scholarly-Primitives-Inspired Research Engine), a multi-agent framework that realizes Scholarly Primitives, a typology of basic humanities scholarship practices, as cooperating roles over a multi-scale close-reading substrate of passages, intra-context graph communities, and cross-context semantic clusters. On a benchmark of peer-reviewed papers in classical Chinese and Greco-Roman Latin scholarship, SPIRE recovers substantially more cited primary-source evidence and receives higher blind human and LLM ratings on four scholarly quality dimensions than LLM, RAG, and generic agentic baselines. Retrieval-tier and agent ablations, with retrieval-volume and answer-length controls, trace its advantage to targeted retrieval, structured evidence selection, and claim-evidence binding. Code and benchmark data are released at https://github.com/YatingPan/SPIRE.

多智能体人文学术证据推理文献挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。