arXiv:2608.19206cs.CLcs.AI2026-08

让大模型的胡说八道变成可验证的科学假说,靠多智能体协作约束

Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses

论文配图:Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
图 1 · 摘自论文原文
  • 用生成与评估双智能体形成思想摩擦,通过语义瓶颈减少重复和噪声
  • 在物理与社会科学领域生成多样且可评估的假说,实验验证了架构有效性
  • 适合需要突破性创意但需经得起实证检验的研究者,尤其在强约束场景下优势明显

当前大语言模型倾向于抑制幻觉,强调事实检索而非创造性组合。本文提出基于Rust的多智能体架构,以叙事想象与执行控制的对比为灵感,构建高熵生成智能体与基于网络的评估智能体之间的认知摩擦。通过低熵语义瓶颈减少噪声与重复。初步实验在物理与社会科学领域生成了多样化、可评估的假说。探索性对照实验对比了完整系统与直接提示、自我反思、移除语义过滤、移除搜索依赖、移除横向视角等变体。结果表明直接提示表现最差,但完整系统未普遍优于简单自我反思。各架构在原创性、可行性、多样性与实证基础间权衡不同,完整系统在面对强物理、实证或制度约束时优势显著。研究不支持孤立幻觉有用,而是强调只有在架构约束、实证基础与显式评估下,推测生成才具价值。

原文摘要 · Abstract (English)

Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also restrict speculative Research and Development (R&D) by encouraging what this work operationally treats as semantic overfitting and diversity collapse. In this paper, we propose a Rust-based multi-agent orchestration that uses the contrast between narrative daydreaming and executive control as a functional analogy, not as a neurocognitive claim. The system instigates an Epistemological Friction loop between a high-entropy generating agent and a web-grounded evaluating agent, mediated by a low-entropy semantic bottleneck intended to reduce noise and repetition. Initial experiments generated diverse, viability-rated hypotheses across physical and social-science domains. We additionally report an exploratory paired baseline and ablation study comparing the full system against direct prompting, self-reflection, removal of the semantic filter, removal of search grounding, and removal of lateral lenses. The results place direct prompting among the weakest conditions across most observed metrics, but they do not show a general superiority of the full system over simple self-reflection. Instead, they suggest that each architecture shifts the balance between originality, feasibility, diversity, and empirical grounding in different ways, and that the full system provides its main advantages when hypotheses must survive strong physical, empirical, or institutional constraints. These findings do not show that hallucination is useful in isolation; they suggest that speculative generation gains value only when constrained by architecture, empirical grounding, and explicit evaluation.

大模型科学假说多智能体幻觉利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。