arXiv:2607.25576cs.AI2026-07

无需系统矩阵,用注意力机制实现快速光声成像重建。

Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention

论文配图:Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention
图 1 · 摘自论文原文
  • 将传感器时序信号作为令牌,直接映射到图像,跳过传统系统矩阵计算。
  • 在46个测试样本上达到0.522的SSIM和22.09 dB的PSNR,显著优于现有方法。
  • 适合需要实时成像的临床场景,重建速度提升至少十倍。

光声断层成像(PAT)结合了生物组织的光学吸收对比度与超声的空间分辨率,但从稀疏视角传感器测量中恢复初始压力分布仍是病态逆问题。迭代压缩感知求解器和展开式深度网络均依赖推理时的系统矩阵,导致临床实时重建计算成本高昂。本文提出传感器注意力网络(SAN),一种基于Transformer的架构,将每个传感器的完整时间序列视为一个令牌,直接从原始测量映射到重建图像,推理时不调用系统矩阵。为训练与基准测试,构建了一个解析的k空间H矩阵,并在匹配几何下通过k-Wave伪谱求解器验证,实现每传感器平均皮尔逊相关系数0.919 ± 0.049;k空间加权与高斯时域衰减协同作用,使能量归一化误差降低49%。在488个增强样本上以血管加权损失训练,于46个保留样本上评估,相比ISTA、split-Bregman总变差(SBTV)和学习型ISTA(LISTA),SAN取得最高均值SSIM(0.522)、PSNR(22.09 dB)和最低均值NMSE(0.233)。配对t检验与威尔科克森符号秩检验表明,SAN在PSNR、NMSE和皮尔逊相关性上显著优于LISTA(p < 1e-8),在所有保真度指标上也优于ISTA和SBTV。通过推理时绕过H矩阵,SAN将重建时间减少至少一个数量级,支持实时PAT重建。

原文摘要 · Abstract (English)

Photoacoustic tomography (PAT) combines the optical absorption contrast of biological tissue with the spatial resolution of ultrasound, yet recovering the initial pressure distribution from sparse-view sensor measurements remains an ill-posed inverse problem. Iterative compressive-sensing solvers and unrolled deep networks both retain a dependence on the system matrix at inference, which leaves real-time clinical reconstruction computationally expensive. This paper proposes the Sensor Attention Network (SAN), a Transformer-based architecture that treats the full time series of each sensor as a token and maps raw measurements directly to the reconstructed image without invoking the system matrix at inference. For training and benchmarking, an analytical k-space H-matrix is constructed and validated against the k-Wave pseudo-spectral solver under matched geometry, achieving a mean per-sensor Pearson correlation of 0.919 +/- 0.049, with k-space apodization and Gaussian temporal damping acting synergistically to reduce the energy-normalized mismatch by 49%. Trained with a vessel-weighted loss on 488 augmented samples and evaluated on 46 held-out samples against ISTA, split-Bregman total variation (SBTV), and learned ISTA (LISTA), SAN attains the highest mean SSIM (0.522) and PSNR (22.09 dB) and the lowest NMSE (0.233). Paired t-tests and Wilcoxon signed-rank tests confirm the superiority of SAN over LISTA on PSNR, NMSE, and Pearson correlation at p < 1e-8, and over ISTA and SBTV on all fidelity metrics. By bypassing the H-matrix at inference, SAN reduces reconstruction time by at least an order of magnitude, supporting real-time PAT reconstruction.

光声成像Transformer实时重建无矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。