arXiv:2603.20729cs.CVcs.AI2026-03

用深度注意力融合图像与井数据,实现无标注的井壁图像分割。

Weakly supervised multimodal segmentation of acoustic borehole images with depth-aware cross-attention

  • 基于阈值伪标签,通过深度感知交叉注意力迭代优化。
  • 融合策略中深度注意力提升准确率,优于直接拼接或普通融合。
  • 适合缺乏标注但有井数据的地质勘探场景使用。

声波井壁图像提供高分辨率井壁结构信息,但大规模解析困难,因密集专家标注稀缺且地下信息本质上多模态。本文提出一种弱监督多模态分割框架,通过学习模型对阈值引导的伪标签进行优化,保留传统阈值与聚类流程的无标注特性,同时引入去噪、置信度感知伪监督和物理结构化融合。实验表明,阈值引导的学习精炼相比原始阈值、去噪阈值和潜在聚类基线具有最强鲁棒性改进。多模态性能高度依赖融合策略:直接拼接增益有限,而深度感知交叉注意力、门控融合与置信度调制显著提升与弱监督参考的一致性。最优模型——置信度门控深度感知交叉注意力(CG-DCA)持续超越基于阈值、仅图像或多模态先前基线。针对性消融分析显示其优势源于置信度感知融合与结构化局部深度交互,而非单纯模型复杂度。跨井分析证实该性能广泛稳定。结果建立了一种实用、可扩展的无标注分割框架,表明当辅助测井数据选择性地结合深度感知信息时,多模态提升最大化。

原文摘要 · Abstract (English)

Acoustic borehole images provide high-resolution borehole-wall structure, but large-scale interpretation remains difficult because dense expert annotations are rarely available and subsurface information is intrinsically multimodal. The challenge is developing weakly supervised methods combining two-dimensional image texture with depth-aligned one-dimensional well-logs. Here, we introduce a weakly supervised multimodal segmentation framework that refines threshold-guided pseudo-labels through learned models. This preserves the annotation-free character of classical thresholding and clustering workflows while extending them with denoising, confidence-aware pseudo-supervision, and physically structured fusion. We establish that threshold-guided learned refinement provides the most robust improvement over raw thresholding, denoised thresholding, and latent clustering baselines. Multimodal performance depends strongly on fusion strategy: direct concatenation provides limited gains, whereas depth-aware cross-attention, gated fusion, and confidence-aware modulation substantially improve agreement with the weak supervisory reference. The strongest model, confidence-gated depth-aware cross-attention (CG-DCA), consistently outperforms threshold-based, image-only, and earlier multimodal baselines. Targeted ablations show its advantage depends specifically on confidence-aware fusion and structured local depth interaction rather than model complexity alone. Cross-well analyses confirm this performance is broadly stable. These results establish a practical, scalable framework for annotation-free segmentation, showing multimodal improvement is maximized when auxiliary logs are incorporated selectively and depth-aware.

多模态分割弱监督地质图像深度注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。