ECHO通过单步块扩散实现快速胸片报告生成,显著提速且保持临床准确。
ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion

- 采用新型直接条件蒸馏框架,解决单步去噪中词元依赖缺失问题。
- 相比自回归模型提升64.33%和60.58%的报告质量指标,推理快8倍。
- 适合需要低延迟、高效率胸片报告生成的临床场景或系统集成。
胸片报告生成(CXR-RG)有望大幅减轻放射科医生的工作负担。然而,传统自回归视觉-语言模型因逐词解码导致推理延迟高。基于扩散的模型虽可通过并行生成降低延迟,但仍需多步去噪迭代。将多步去噪压缩至单步可进一步提速,但常因词元因子化去噪器引入的均值场偏差导致文本连贯性下降。为此,我们提出ECHO,一种高效的扩散型视觉-语言模型(dVLM),用于胸片报告生成。ECHO通过新颖的直接条件蒸馏(DCD)框架实现稳定的一步一区块推理,该框架利用在线策略扩散轨迹构建非因子化监督,以捕捉词元间的联合依赖关系。此外,我们引入响应不对称扩散(RAD)训练策略,在不损失模型效果的前提下提升训练效率。大量实验表明,ECHO超越当前最优自回归方法,在RaTE和SemScore上分别提升64.33%和60.58%,同时实现高达8倍的推理速度提升,临床准确性几乎无损。
原文摘要 · Abstract (English)
Chest X-ray report generation (CXR-RG) has the potential to substantially alleviate radiologists' workload. However, conventional autoregressive vision--language models (VLMs) suffer from high inference latency due to sequential token decoding. Diffusion-based models offer a promising alternative through parallel generation, but they still require multiple denoising iterations. Compressing multi-step denoising to a single step could further reduce latency, but often degrades textual coherence due to the mean-field bias introduced by token-factorized denoisers. To address this challenge, we propose \textbf{ECHO}, an efficient diffusion-based VLM (dVLM) for chest X-ray report generation. ECHO enables stable one-step-per-block inference via a novel Direct Conditional Distillation (DCD) framework, which mitigates the mean-field limitation by constructing unfactorized supervision from on-policy diffusion trajectories to encode joint token dependencies. In addition, we introduce a Response-Asymmetric Diffusion (RAD) training strategy that further improves training efficiency while maintaining model effectiveness. Extensive experiments demonstrate that ECHO surpasses state-of-the-art autoregressive methods, improving RaTE and SemScore by \textbf{64.33\%} and \textbf{60.58\%} respectively, while achieving up to \textbf{$8\times$} inference speedup with negligible degradation in clinical accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。