arXiv:2608.17501cs.AIcs.LG2026-08

用本地大模型发现可追溯的研究问题,避免依赖封闭模型的幻觉和偏见。

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

论文配图:SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
图 1 · 摘自论文原文
  • 基于文献库构建证据图,自动识别研究空白
  • 生成的问题包含假设、目标与未解歧义,可审计可追踪
  • 全程本地运行,不调用外部闭源模型,保障隐私

当前全自动AI科学家在提出研究问题时,高度依赖封闭的前沿语言模型,其推理过程受黑箱参数知识和自我引导的文献检索影响,难以验证真实性,易产生幻觉与偏见。若将敏感研究数据传至外部API,还存在隐私与数据治理风险。本文提出结构缺口假设代理(SGHA),一个完全本地运行的、以文献语料为先的研究问题发现系统。它将科学文献库组织为带证据关联的论文对象与类型化证据图,检测跨论文的未解结构模式,在问题形成前筛选候选缺口,并输出可追溯的研究问题族。系统能明确列出假设、目标、成功标准及剩余模糊点。所有语言模型组件均基于本地部署的90亿参数开源模型运行,无需调用闭源前沿模型接口。在五个机器学习领域对比AI Scientist-v2的问题生成模块,结果表明:显式文献结构与证据约束推理,可在不依赖前沿模型的情况下实现可检查、有前景的研究问题生成。

原文摘要 · Abstract (English)

Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model-specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data-governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus-first research-problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research-problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM-based components of SGHA are executed using a locally served open-weight 9B language model, without requiring proprietary frontier-model APIs. We compare SGHA with the AI Scientist-v2 idea formulation module in five machine-learning domains. Our results suggest that explicit corpus structure and evidence-constrained reasoning can support promising, inspectable research-problem formulation without relying on frontier models during generation or verification.

研究问题发现本地大模型可解释性证据图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。