arXiv:2604.19512eess.IV2026-04

为超声图像重建开发了更贴近临床的评估标准。

Defining Robust Ultrasound Quality Metrics via an Ultrasound Foundation Model

论文配图:Defining Robust Ultrasound Quality Metrics via an Ultrasound Foundation Model
图 1 · 摘自论文原文
  • 基于超声基础模型构建感知距离与无参考质量评分
  • 能准确反映分割任务性能下降,优于传统指标
  • 适合临床医生偏好预测,助力超声图像优化

临床医生缺乏量化超声重建诊断价值的系统框架。现有标准如PSNR和VGG-LPIPS无法体现超声特有的物理特性或声学成像结构细节。本文提出基于TinyUSFM的评估体系,包含两个指标:TinyUSFM-uLPIPS(基于多层令牌关系的全参考感知距离)和TinyUSFM-NRQ(利用干净流形建模与最差区域聚合的可部署无参考质量分数),用于检测局部有害伪影。实验表明,该框架具备四大优势:1)任务相关性——TinyUSFM-uLPIPS在语义任务损伤下表现更优,能准确反映分割Dice分数下降,而VGG类指标失效;2)跨器官可比性——在不同解剖部位及域偏移数据上保持稳定评分尺度与一致严重度排序;3)与PSNR一致的敏感性——无需真实图像即可提供可靠质量评分,且与传统保真度基准(如PSNR)一致;4)临床实用性——专家偏好预测准确率从47.2%提升至72.8%,并生成超声科医师更青睐的超分辨率重建结果。通过整合这些优势,本工作建立了一个模态对齐的标准,真正弥合算法性能与诊断效用之间的鸿沟。代码已开源:https://github.com/sextant-fable/US-Metrics。

原文摘要 · Abstract (English)

Clinicians lack a principled framework to quantify diagnostic utility in ultrasound reconstructions. Existing standards like PSNR and VGG-LPIPS are inadequate, failing to account for modality-specific physics or the structural nuances of acoustic imaging. We close this gap with a TinyUSFM-based evaluation framework featuring two distinct metrics: TinyUSFM-uLPIPS, a full-reference perceptual distance based on multi-layer token relations, and TinyUSFM-NRQ, a deployable no-reference quality score utilizing clean-manifold modeling and worst-region aggregation to detect localized harmful artifacts. We demonstrate that the presented metrics have four unique advantages: 1) Task-linked quality, where TinyUSFM-uLPIPS achieves superior calibration with semantic task damage, accurately reflecting Dice-score drops in segmentation where VGG-based metrics fail; 2) Cross-organ comparability, maintaining stable scoring scales and consistent severity rankings across diverse anatomical sites and domain-shifted data; 3) PSNR-consistent sensitivity, with TinyUSFM-NRQ providing a reliable quality score without ground-truth images that remains consistent with traditional fidelity benchmarks (i.e. PSNR); and 4) Clinical utility, improving the prediction of expert preference from 47.2$\%$ to 72.8$\%$ accuracy and producing super-resolution reconstructions preferred by sonographers. By integrating these advantages into a unified assessment and optimization loop, this work establishes a modality-aligned standard that finally bridges the gap between algorithmic performance and diagnostic utility. Our code is available at https://github.com/sextant-fable/US-Metrics.

超声评估质量度量基础模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。