arXiv:2608.24011cs.CLcs.AI2026-08

让AI读懂古文不再靠猜,而是像侦探一样找证据、验结论。

SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding

论文配图:SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding
图 1 · 摘自论文原文
  • 用多个专业AI agent协作,分步推理而非直接回答。
  • 在AncientDoc数据集上超越更大模型,关键指标提升12.3%。
  • 适合需要高可信度解读的古籍研究与数字人文项目。

中文古籍理解需融合视觉、语言与历史推理。现有大视觉语言模型多采用单次生成的黑箱模式,常产生过度自信且缺乏依据的回答。为此,我们提出SAGE——一种基于证据的多智能体框架,将古籍理解重构为证据支撑的推理过程。SAGE通过任务规划、工具辅助获取证据、逐项验证主张以及受限状态下的动态重规划,实现有限证据搜索、答案修正与必要时拒答。在AncientDoc基准测试中,SAGE在三种LVLM骨干网络下均优于直接回答基线。尤为突出的是,使用Qwen3.5-9B的SAGE在多数指标上超越更大规模的单体模型,凸显结构化证据推理对提升性能的重要性,远超单纯模型扩展。

原文摘要 · Abstract (English)

Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning. Current Large Vision-Language Models (LVLMs) typically rely on an opaque, single-pass generation paradigm, often producing overconfident and weakly grounded responses. To address this, we propose SAGE, an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation. SAGE coordinates specialized agents for task-aware planning, tool-mediated evidence acquisition, claim-level verification, and bounded replanning under a constrained shared-state runtime. This design supports bounded evidence seeking, answer revision, and abstention when grounding is insufficient. Experiments on the AncientDoc benchmark show that SAGE consistently outperforms matched direct-answering baselines across three LVLM backbones. Remarkably, SAGE with Qwen3.5-9B surpasses much larger monolithic LVLMs on most evaluated metrics, highlighting the importance of structured, evidence-grounded inference beyond model scaling.

古籍理解多智能体证据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。