构建临床相关合成数据集,评估深度学习在光声成像中的重建效果。
Benchmarking Deep Learning-Based Reconstruction Methods for Photoacoustic Computed Tomography with Clinically Relevant Synthetic Datasets
- 用1.1万+个真实乳腺结构的合成数据集评估图像重建方法。
- 深度学习模型虽在传统指标上表现好,但常无法准确恢复病灶。
- 适合研究光声成像算法或需要临床相关评估的团队使用。
近年来,基于深度学习(DL)的光声计算机断层成像(PACT)图像重建方法发展迅速。然而,大多数现有方法未采用标准化数据集,其评估依赖的传统图像质量(IQ)指标可能缺乏临床意义。缺乏标准化的临床相关评估框架,阻碍了公平比较,并引发对报告进展可复现性和可靠性的担忧。本文提出一个基准测试框架,提供开源、解剖学合理的合成数据集和评估策略,用于评估PACT中基于DL的声学反演方法。数据集包含超过11,000个二维随机乳腺对象,包含临床相关的病灶,以及在不同建模复杂度下的配对测量数据。评估策略结合传统与任务导向的图像质量度量,以评估重建保真度和临床实用性。初步基准测试展示了该框架的有效性,通过对比基于DL与基于物理的重建方法,发现尽管某些基于DL的方法在传统指标上表现良好,但在病灶恢复方面存在明显缺陷。这凸显了传统指标的局限性,强调了任务导向评估的重要性。该框架支持对2D PACT中基于DL的声学反演方法进行系统性比较,通过整合临床相关合成数据与严格评估协议,实现可复现、客观的性能评估,促进方法开发与系统优化。
原文摘要 · Abstract (English)
Deep learning (DL)-based image reconstruction methods for photoacoustic computed tomography (PACT) have developed rapidly in recent years. However, most existing methods have not employed standardized datasets, and their evaluations rely on traditional image quality (IQ) metrics that may lack clinical relevance. The absence of a standardized framework for clinically meaningful IQ assessment hinders fair comparison and raises concerns about the reproducibility and reliability of reported advancements in PACT. A benchmarking framework is proposed that provides open-source, anatomically plausible synthetic datasets and evaluation strategies for DL-based acoustic inversion methods in PACT. The datasets each include over 11,000 two-dimensional (2D) stochastic breast objects with clinically relevant lesions and paired measurements at varying modeling complexity. The evaluation strategies incorporate both traditional and task-based IQ measures to assess fidelity and clinical utility. A preliminary benchmarking study is conducted to demonstrate the framework's utility by comparing DL-based and physics-based reconstruction methods. The benchmarking study demonstrated that the proposed framework enabled comprehensive, quantitative comparisons of reconstruction performance and revealed important limitations in certain DL-based methods. Although they performed well according to traditional IQ measures, they often failed to accurately recover lesions. This highlights the inadequacy of traditional metrics and motivates the need for task-based assessments. The proposed benchmarking framework enables systematic comparisons of DL-based acoustic inversion methods for 2D PACT. By integrating clinically relevant synthetic datasets with rigorous evaluation protocols, it enables reproducible, objective assessments and facilitates method development and system optimization in PACT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。