用状态机控制生成流程,让科研助手不胡说、有依据、可追溯。
Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval
- 通过状态机分步控制:相关性→置信度→知识,避免随意回答。
- 专家评测中更受青睐,能准确处理边界条件并提供去重引用。
- 适合高要求科研场景,支持光电等领域的可扩展知识库构建。
大语言模型虽加速文献整合,但易产生幻觉和错误引用,限制其在专业工作流中的应用。我们提出 RA-FSM(研究助理-有限状态机),一种基于 GPT 的模块化科研助手,采用有限状态控制循环:相关性 → 置信度 → 知识。系统依托向量检索与确定性引文流程,控制器可过滤无关问题、评估可答性、拆解问题,并仅在必要时触发检索,输出带置信标签及原文内去重引用的答案。采用分级摄入流程,从期刊、会议、索引、预印本和专利构建领域知识库,同时写入密集向量索引与结构化指标存储。我们在光子学领域对六类任务(分析推理、数值分析、方法批判、对比综述、事实提取、应用设计)进行评估。在盲测 A/B 对比中,领域专家更偏好 RA-FSM,优于强基线 Notebook LM(NLM)与单次调用的默认 GPT API,认为其边界条件处理更强、证据使用更可信。覆盖率与新颖性分析显示,RA-FSM 在超越 NLM 的同时,可调节延迟与成本开销。该设计强调透明、有据的回答,适用于高风险技术工作,且可推广至其他科学领域。
原文摘要 · Abstract (English)
Large language models accelerate literature synthesis but can hallucinate and mis-cite, limiting their usefulness in expert workflows. We present RA-FSM (Research Assistant - Finite State Machine), a modular GPT-based research assistant that wraps generation in a finite-state control loop: Relevance -> Confidence -> Knowledge. The system is grounded in vector retrieval and a deterministic citation pipeline. The controller filters out-of-scope queries, scores answerability, decomposes questions, and triggers retrieval only when needed, and emits answers with confidence labels and in-corpus, de-duplicated references. A ranked-tier ingestion workflow constructs a domain knowledge base from journals, conferences, indices, preprints, and patents, writing both to a dense vector index and to a relational store of normalized metrics. We implement the system for photonics and evaluate it on six task categories: analytical reasoning, numerical analysis, methodological critique, comparative synthesis, factual extraction, and application design. In blinded A/B reviews, domain experts prefer RA-FSM to both a strong Notebook LM (NLM) and a vanilla Default GPT API call single-pass baseline, citing stronger boundary-condition handling and more defensible evidence use. Coverage and novelty analyses indicate that RA-FSM explores beyond the NLM while incurring tunable latency and cost overheads. The design emphasizes transparent, well-cited answers for high-stakes technical work and is generalizable to other scientific domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。