arXiv:2605.11154astro-ph.IMcs.AI2026-05中稿 · publication in PAS…

用大模型和信息论量化天体方法的可复现性,发现文字描述难还原真实代码。

Quantifying the Reconstructability of Astrophysical Methods with Large Language Models and Information Theory: A Case Study in Spectral Reconstruction

  • 将算法复现视为大模型生成的概率分布,用熵与散度衡量文本对实现的约束力。
  • 在稀疏光度数据重建中,文本增加虽清晰化结构,但实现差异仍存,存在不可消除的熵底。
  • 揭示了专家隐性知识缺失导致大模型无法精准复现科学校准,适合关注可复现性的研究者。

现代天体物理研究依赖复杂数据分析流程,但发表描述常缺乏可计算复现所需细节。本文提出信息论框架,量化从文本描述复现方法的有效性。将算法复现建模为大语言模型(LLMs)生成的概率分布,利用香农熵与Jensen-Shannon散度测量文本对有效实现空间的约束程度。以海王星外天体(TNO)光谱重建为例,通过不同文本层级(标题、摘要、方法)提示前沿大模型,发现文本增多虽提升算法结构清晰度,但实现层面的方差无法消除,存在‘熵底’——即多个不同实现仍符合显式指令。进一步将重构算法转化为可执行流程,结果显示大模型虽能恢复核心功能,却系统性遗漏严格科学校准所需的隐性专业知识。本研究证明大模型可作为零样本诊断工具,审计方法透明度,帮助作者识别缺失的结构约束,维护自动化研究时代的科学完整性。

原文摘要 · Abstract (English)

Modern astrophysical studies rely heavily on complex data analysis pipelines; however, published descriptions often lack the detail required for computational reproducibility. In this work, we present an information-theoretic framework to quantify how effectively a method can be reconstructed from its written description. By treating algorithmic reconstruction as a probability distribution generated by Large Language Models (LLMs), we utilize Shannon entropy and Jensen-Shannon divergence to measure how strongly text constrains the hypothesis space of valid implementations. We demonstrate this approach through a case study of Trans-Neptunian Object (TNO) spectral reconstruction from sparse photometry. By prompting frontier LLMs with varying levels of manuscript text (Title, Abstract, and Methods), we find that while increasing text successfully clarifies the overall algorithmic structure, it fails to eliminate variance at the implementation level. This persistent variance establishes an "entropy floor," demonstrating that multiple divergent implementations remain consistent with explicit instructions. To evaluate practical reproducibility, we convert these reconstructed algorithms into executable pipelines. Our results reveal that, while LLMs easily recover core functional methodologies, they systematically fail to infer the tacit expert knowledge required for strict scientific calibration. This pilot study demonstrates that LLMs can be repurposed as a zero-shot diagnostic tool to audit methodological transparency, helping authors identify missing structural constraints and preserve scientific integrity in an era of automated research.

大模型可复现性天体物理信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。