arXiv:2605.22300cs.AIcs.LG2026-05

跨领域实验揭示协作AI在科学推断中的真实价值边界

Cross-domain benchmarks reveal when coordinated AI agents improve scientific inference from partial evidence

论文配图:Cross-domain benchmarks reveal when coordinated AI agents improve scientific inference from partial evidence
图 1 · 摘自论文原文
  • 构建四类跨学科任务基准,对比协作AI与单通道方法
  • 气候-媒介疾病预测达AUROC 0.944,系外行星验证达0.955
  • 仅当性能、可追溯性或表征有实证提升时,协作才真正增益

科学证据常分散于不同仪器、数据库和学科,单一来源无法完整记录现象。本文通过涵盖四大任务的跨领域基准评估:分子结构转音乐表示、科学范式变迁检测、虫媒疾病暴发识别、系外行星候选验证。每项任务均采用冻结评估组、预设评分协议、明确基线、消融实验或零模型对照及局限性声明。结果定义三种运行模式:当各学科仅捕捉现象部分时,跨通道组合优于单通道基线(气候-媒介疾病:AUROC 0.944;系外行星验证:AUROC 0.955);但系外行星工作流与强综合摘要基线无显著差异,表明分解不总提升表现;当某信号主导时(如范式变迁),协调主要提升解释性与可追溯性;分子声学转换中,收益为表征而非预测。ScienceClaw x Infinite 提供可审计的成果与溯源层。因此,仅当性能、溯源或表征主张被显式比较支持时,协调才有价值。

原文摘要 · Abstract (English)

Scientific evidence often spans instruments, databases, and disciplines, so no single source records the full phenomenon. This makes it difficult to determine when coordinated AI agents add value over simpler scientific workflows. We evaluate this question with a cross-domain benchmark spanning four scientific tasks: mapping molecular structure into musical representations, detecting historical paradigm shifts in science, identifying vector-borne disease emergence, and vetting transiting-exoplanet candidates. Each case uses a frozen evaluation panel, predefined scoring protocols, explicit baselines, ablations or null controls, and stated limitations. The results define three operating regimes. When different disciplines each capture only part of the phenomenon, cross-channel composites improve over single-channel baselines: climate-vector emergence reaches AUROC 0.944 and exoplanet vetting reaches AUROC 0.955. However, the exoplanet workflow is effectively tied with a strong combined-summary baseline, showing that decomposition does not always improve top-line performance. When one signal dominates, as in paradigm-shift detection, coordination mainly improves interpretation and traceability. For molecular sonification, the gain is representational rather than predictive. ScienceClaw x Infinite provides the auditable artifact and provenance layer for this evaluation. The benchmark therefore assigns value to coordination only when the corresponding performance, provenance, or representation claim is supported by explicit comparators.

AI科学推理跨领域基准可追溯性多智能体协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。