arXiv:2601.03436astro-ph.IMcs.AI2026-01

用大模型打造可信赖的科研助手,专攻引力波领域技术问答

MARVEL: A Multi Agent-based Research Validator and Enabler using Large Language Models

  • 分快速查询与深度搜索双模式,结合检索增强与树搜索策略
  • 在探测器运维问答上超越GPT-4o mini,文献问答持平
  • 适合需要精准引用、私有网络运行的科研团队使用

我们提出MARVEL(https://ligogpt.mit.edu/marvel),一个可本地部署的开源框架,用于领域感知的问答与科研辅助。针对科学团队对数字助手日益增长的需求——能理解高度技术性数据、精准引用并运行于受认证网络中,MARVEL提供快速路径和更严谨的DeepSearch模式。该模式融合检索增强生成与蒙特卡洛树搜索,探索互补子查询,为有潜力分支分配更多算力,并维护全局证据记录以保留来源。系统基于精心构建的语义索引,涵盖研究论文、博士论文、LIGO文档及长期探测器电子日志,必要时辅以定向网络搜索。由于无法在私有数据上直接对比商用大模型,我们在两个公开的替代数据集上评估了MARVEL,其在文献类查询上达到GPT-4o mini水平,在探测器运行内容上显著优于后者,凸显领域检索与引导推理的重要性。通过开放完整框架与评估数据集,旨在为构建领域专用科研助手提供可复现基础。

原文摘要 · Abstract (English)

We present MARVEL (https://ligogpt.mit.edu/marvel), a locally deployable, open-source framework for domain-aware question answering and assisted scientific research. It is designed to address the increasing demands of a digital assistant for scientific groups that can read highly technical data, cite precisely, and operate within authenticated networks. MARVEL combines a fast path for straightforward queries with a more deliberate DeepSearch mode that integrates retrieval-augmented generation and Monte Carlo Tree Search. It explores complementary subqueries, allocates more compute to promising branches, and maintains a global evidence ledger that preserves sources during drafting. We applied this framework in the context of gravitational-wave research related to the Laser Interferometer Gravitational-wave Observatory. Answers are grounded in a curated semantic index of research literature, doctoral theses, LIGO documents, and long-running detector electronic logbooks, with targeted web searches when appropriate. Because direct benchmarking against commercial LLMs cannot be performed on private data, we evaluated MARVEL on two publicly available surrogate datasets that capture comparable semantic and technical characteristics. On these benchmarks, MARVEL matches a GPT-4o mini baseline on literature-centric queries and substantially outperforms it on detector-operations content, where domain retrieval and guided reasoning are decisive. By making the complete framework and evaluation datasets openly available, we aim to provide a reproducible foundation for developing domain-specific scientific assistants.

科研助手大模型引力波问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。