arXiv:2508.02789cs.AI2025-08

让AI像科学家一样自我调整推理过程,提升科学问题解答的准确性和可控性。

Cognitive Loop via In-Situ Optimization: Self-Adaptive Reasoning for Science

  • 通过原位优化构建认知循环,让大模型自主调整推理策略
  • 在生物医学题库上达22.37%准确率,比基线提升161.64%
  • 开放设计可观察不确定性与推理路径,适合科研人员深度参与

人工智能在动态条件下形成、演化并测试新型思维模式的能力,是科学发现中高级认知的关键。现有AI系统分为两类:一类基于非推理模型,嵌入对人类思维的预设假设;另一类为推理模型,但将推理直觉抽象化,使用户难以掌控。为提升科学家使用AI进行科学探索的效率,需兼具准确性、透明性与可操控性。为此,本文提出一种新方法——原位优化的认知循环(CLIO),使大语言模型能够自主制定解题策略,在自信心不足时自我调整,并最终输出可信结论。通过开放设计,科学家可追踪不确定性水平、理解基于图结构的信念生成过程,并主动干预修正。无需额外微调,GPT-4.1搭配CLIO在人类最后考试(HLE)文本生物医学题库上取得22.37%准确率,相比基线模型提升13.82个百分点(相对提升161.64%),超越OpenAI o3在高/低推理努力模式下的表现。研究还发现,内部不确定性波动是决定结果准确性的关键因素,揭示了该框架在科学决策中的可观测性与控制力。

原文摘要 · Abstract (English)

The capacity for artificial intelligence (AI) to formulate, evolve, and test altered thought patterns under dynamic conditions indicates advanced cognition that is crucial for scientific discovery. The existing AI development landscape falls into two categories: 1) frameworks over non-reasoning models that natively incorporate opinions on how humans think, and 2) reasoning models that abstract precise control of the reasoning intuition away from end users. While powerful, for scientists to maximize utility of AI in scientific discovery, they not only require accuracy and transparency in reasoning, but also steerability. Hence, we introduce an alternative approach that enables deep and precise control over the reasoning process called: a cognitive loop via in-situ optimization (CLIO). CLIO enables large language models (LLMs) to self-formulate ways of approaching a problem, adapt behavior when self-confidence is low, and ultimately provide scientists with a final belief or answer. Through CLIO's open design, scientists can observe uncertainty levels, understand how final belief states are formulated using graph structures, and interject corrections. Without any further post-training, OpenAI's GPT-4.1 with CLIO yields an accuracy of 22.37\% in text-based biology and medicine questions on Humanity's Last Exam (HLE). This yields a 13.82\% net or 161.64\% relative increase when compared to the base GPT-4.1 model and surpasses OpenAI's o3 performance in high and low reasoning effort modes. We further discovered that oscillations within internal uncertainty measures are key in determining the accuracy of CLIO's results, revealing how its open design and internal mechanisms can provide insight and control into scientific decision-making processes.

认知循环科学推理大模型控制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。