arXiv:2607.22824physics.med-phcs.AI2026-07

用智能体自动完成CT重建方法的测试与优化,发现理想数据排名不等于真实噪声下的表现。

Agentic Autoresearch for CT Reconstruction

论文配图:Agentic Autoresearch for CT Reconstruction
图 1 · 摘自论文原文
  • 构建智能体循环,自动调参、运行并评估26种重建方法。
  • 在真实噪声下,顶尖方法排名逆转,理想数据榜单失效。
  • 强调多因素混合挑战才能衡量模型真正泛化能力,适合医学影像研究者。

公平比较CT重建方法耗时且高度依赖人工,许多基准使用理想化数据。本文探讨大型语言模型(LLM)智能体能否自主完成重建研究,并验证理想数据上的排名是否能预测真实噪声下的表现。我们构建了一个智能体闭环:智能体修改求解器,执行短时集群任务,读取一个固定指标(视场内相对于FBP基线的校准余量分数),并据此修正。所有方法共享同一可微分的扇形束投影器。我们在梅奥低剂量CT(噪声受限)和无噪声DL-Sparse-View挑战中的128视角乳腺任务上对26种方法进行基准测试,使用验证集选定迭代次数并在保留测试集上评分。每个训练好的乳腺模型均在噪声输入(I_0 = 10^5光子)下重新评分(无需重训练),并另在匹配噪声下重新训练。智能体独立实现、调优与评估全部26种方法,并将它们组合为一个仅969参数的紧凑求解器,在梅奥数据上以1%水平与冠军相当,仅用其0.4%参数。基准测试揭示一组统计上不可区分的顶级方法,而非单一胜者。轻微输入噪声几乎反转乳腺任务排名:无噪声冠军(监督图像去噪器,头围分数0.89)降至0.00,而学习型原对偶方法跃升至冠军(0.72至0.93)。因此,理想数据排行榜无法预测鲁棒性。该反转是迁移效应,非永久缺陷:在匹配噪声下重训练后,清洁排名显著恢复(斯皮尔曼等级相关系数0.04提升至0.61)。噪声只是开放问题中一个最简单的混杂因子(如束硬化、散射、解剖结构、疾病),故单因素挑战无法认证通用性。基准应同时模拟广泛现实因素。

原文摘要 · Abstract (English)

Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large language model (LLM) agent can do the labor of reconstruction research on its own, and whether a ranking measured on ideal data predicts behavior under realistic noise. We built an agentic loop: the agent edits a solver, runs a short cluster job, reads one frozen metric, and revises. The metric is a calibrated headroom score against the FBP baseline, inside the field of view; every method shares the same differentiable fan-beam projector. We benchmarked 26 methods on Mayo low-dose CT (noise-limited) and a 128-view sparse-view breast task from the noiseless DL-Sparse-View Challenge, with validation-selected iterations scored on a held-out test set. Every trained breast model was then re-scored on noisy inputs (I_0 = 10^5 photons) without retraining, and separately retrained on matched noise. The agent independently implemented, tuned, and benchmarked all 26 methods, and recombined them into a compact solver of 969 parameters that ties the top Mayo tier at the 1% level using 0.4% of the champion's parameters. Benchmarking gives a tier of statistically indistinguishable top methods, not one winner. Mild input noise nearly inverts the breast ranking: the noiseless champion (a supervised image denoiser, hr 0.89) collapses to 0.00, while a learned primal-dual method rises to champion (0.72 to 0.93). An ideal-data leaderboard therefore does not predict robustness. The inversion is a transfer effect, not a permanent deficit: retraining on matched noise restores much of the clean ranking (Spearman rho 0.04 to 0.61). Noise is only the easiest confounder in an open-ended set (beam hardening, scatter, anatomy, disease), so no single-factor challenge certifies generality. Benchmarks should model a broad spectrum of realistic factors at once.

CT重建智能体医学影像鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。