用分层知识库和多智能体写作,让论文报告永远不自相矛盾、可追溯时间点。
Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research
- 建立可信度分级的知识库,统一管理证据与数据,确保信息一致性。
- 生成报告时严格限定时间截点,消除未来数据泄露,零自相矛盾。
- 适合需要高可信度、可复现的科研或金融分析场景。
大语言模型生成的长篇研究报告常出现内容漂移、自我矛盾及溯源缺失问题:同一指标数值不一致,传闻被当作权威数据引用。本文提出双层智能体系统,将动态维护的时空定点知识库与报告生成分离。确定性‘图书管理员’将带时间戳的来源整合进可信度分级的本体结构中,构建包含证据卡、权威指标账本和主张图的实时真相源,而非对原始片段进行每查询检索。便携式多智能体‘写作者’运行时可在任意知识截止时间T生成无矛盾、有依据的报告,仅读取时间≤T的证据(无前瞻);红队验证结果反向回传至图书管理员。评估基于自收集的6,130条公开数据源(涵盖295家发行人、11个行业、美国劳工统计局报告及维基百科),生成555,926张证据卡。单一知识库支撑四份不同论点的时点报告,并完成八次可复现实验,核心指标通过确定性质量控制门控,经缺陷注入元评估验证召回率与精确率均为1.0。共享指标账本消除6,845处跨部分矛盾至零。层级优先选择在22个黄金案例中全对,而流行度优先基线仅正确9/22;可信度分级避免媒体数据泄露,政府统计从不覆盖公司自身披露。红队反驳可反向传播并自动修正后续运行,无需人工干预。重放测试在七个截止点均无前瞻违规,知识库规模由235,373增至555,312张卡片。难度分级模型调度超越全Opus质量上限,且运行速度提升3.7倍。
原文摘要 · Abstract (English)
Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing. A deterministic "librarian" ingests timestamped sources into a trust-tiered ontology, layering evidence cards, an authoritative metric ledger, and a claim graph into an always-current source of truth, not per-query RAG over raw chunks. A portable multi-agent "writer" runtime then composes a contradiction-free, evidence-grounded report at any knowledge cutoff T, reading only evidence with as_of <= T (no look-ahead); red-team verdicts flow back into the librarian. We evaluate on a self-collected, public corpus of 6,130 sources yielding 555,926 evidence cards (SEC EDGAR filings across 295 issuers and 11 sectors, U.S. Bureau of Labor Statistics releases, and Wikipedia). From the one library we compose four point-in-time reports on distinct theses and run eight reproducible experiments, whose headline metrics come from a deterministic quality-control gate, itself validated by defect-injection meta-evaluation at recall 1.0 and precision 1.0. A shared metric ledger removes 6,845 cross-section contradictions to zero. Tier-first selection is correct on 22/22 gold cases where a popularity-first baseline scores only 9/22; trust tiering leaks zero media-sourced numbers, and no government statistic displaces a company's own filing. A red-team refutation propagates back and self-corrects a later run with zero manual edits. Replay exhibits zero look-ahead violations across seven cutoffs while the library grows from 235,373 to 555,312 cards. Difficulty-tiered model routing exceeds the all-Opus quality ceiling while running 3.7x faster than serial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。