通过病变特异性文本与多尺度感知提升肝肿瘤分割精度
TexLiverNet: Leveraging Medical Knowledge and Spatial-Frequency Perception for Enhanced Liver Tumor Segmentation
- 用病变级文本标注增强图文融合,精准捕捉肿瘤细节
- 在公开与私有数据集上均超越现有最优方法
- 适合需要高精度肝肿瘤分割的临床研究与辅助诊断
将文本信息与影像数据融合对提升肝肿瘤分割诊断准确性至关重要。然而,现有医学多模态数据集仅提供通用文本注释,缺乏对病灶的特异性描述,难以提取细微特征,尤其在肿瘤边界和小病灶的细粒度分割方面表现不足。为此,我们构建了包含病变特异性文本注释的肝肿瘤数据集,并提出TexLiverNet模型。该模型采用基于代理的交叉注意力模块,高效融合文本与视觉特征,显著降低计算开销。同时,引入增强的空间感知与自适应频域感知机制,精确勾勒病灶边界,抑制背景干扰,恢复小病灶的精细结构。在公共与私有数据集上的全面评估表明,TexLiverNet在性能上优于当前最先进的方法。
原文摘要 · Abstract (English)
Integrating textual data with imaging in liver tumor segmentation is essential for enhancing diagnostic accuracy. However, current multi-modal medical datasets offer only general text annotations, lacking lesion-specific details critical for extracting nuanced features, especially for fine-grained segmentation of tumor boundaries and small lesions. To address these limitations, we developed datasets with lesion-specific text annotations for liver tumors and introduced the TexLiverNet model. TexLiverNet employs an agent-based cross-attention module that integrates text features efficiently with visual features, significantly reducing computational costs. Additionally, enhanced spatial and adaptive frequency domain perception is proposed to precisely delineate lesion boundaries, reduce background interference, and recover fine details in small lesions. Comprehensive evaluations on public and private datasets demonstrate that TexLiverNet achieves superior performance compared to current state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。