arXiv:2609.05816cs.CV2026-09

针对裂缝分割中漏检与误判问题,提出条件化语义-视觉修正方法。

CoRe-SAM3: Conditional Semantic--Visual Reconciliation for SAM3 Crack Segmentation

论文配图:CoRe-SAM3: Conditional Semantic--Visual Reconciliation for SAM3 Crack Segmentation
图 1 · 摘自论文原文
  • 基于语义与视觉表征的功能差异,设计条件修正机制
  • 裂缝分割平均IoU提升至70.47%,仅增加18.9K参数
  • 适合需要轻量高精度裂缝检测的工程应用

裂缝分割需同时识别目标语义并精确恢复细长、低对比度、拓扑连续的局部结构。尽管SAM3具备强开放概念分割能力,其直接应用于裂缝领域仍存在弱裂缝漏检、激活类裂缝背景区域及局部边界误差。本文在五个裂缝数据集上诊断SAM3内部提示引导语义表示与原始视觉表示的功能差异:语义表示已携带多数任务信息,而视觉表示效用依赖当前语义状态;直接融合二者未带来一致性能提升。基于此,提出条件语义-视觉修正(CoRe)。CoRe以语义预测为主决策路径,通过轻量语义校准调整目标域决策映射,并利用空间对齐的原始视觉证据生成零初始化、有界且正则化的条件残差,选择性修正现有预测。在五个数据域上,CoRe-SAM3将裂缝平均IoU从62.34%提升至70.47%,clDice从81.98%提升至89.24%,仅引入18.914K可训练参数。预测转换分析显示,CoRe平均纠正34.38%的原始错误,对原SAM3正确分类像素的误伤率仅为0.23%。结果表明,基于内部表示功能差异的约束性预测修正,是具有强任务语义先验的视觉基础模型实现高效目标域适配的有效策略。

原文摘要 · Abstract (English)

Crack segmentation requires a model to recognize target semantics while accurately recovering thin, low-contrast, and topologically continuous local structures. Although SAM3 provides strong open-concept segmentation, its direct application to the crack domain still misses weak cracks, activates crack-like background regions, and produces local boundary errors. We first diagnose the functional differences between the internal prompt-conditioned semantic representation and native visual representation of SAM3 on five crack datasets. The results show that the semantic representation already carries most task information for crack prediction, whereas the utility of the visual representation depends on the current semantic state. Directly combining the two representations does not yield consistent gains. Based on this finding, we propose Conditional Semantic--Visual Reconciliation, termed CoRe. CoRe retains semantic prediction as the primary decision path, applies lightweight semantic calibration to adjust the target-domain decision mapping, and uses spatially aligned native visual evidence to generate a zero-initialized, bounded, and regularized conditional residual that selectively corrects existing predictions. Across five domains, CoRe-SAM3 improves the average Crack IoU from 62.34% to 70.47% and clDice from 81.98% to 89.24%, while introducing only 18.914 K trainable parameters. Prediction-transition analysis further shows that CoRe corrects an average of 34.38% of native errors, with a damage rate of only 0.23% on pixels correctly classified by native SAM3. These results demonstrate that constrained prediction correction based on the functional differences between internal representations provides an effective and parameter-efficient target-domain adaptation strategy for vision foundation models with strong task-specific semantic priors.

裂缝分割视觉修正SAM3轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。