针对裂缝分割设计轻量级结构-方向建模,效果超越复杂混合架构。
Rethinking Efficient Crack Segmentation with Task-Aligned Structural-Directional Modeling

- 将裂缝分割视为稀疏结构恢复,聚焦方向连续性与局部证据保留。
- 在4个公开数据集上16项指标均达最优或并列最优,最高精度超基线。
- 模型仅0.47M参数,推理速度快,适合边缘部署,适配低资源场景。
现有裂缝分割方法多采用通用语义分割设计,使用更强的主干网络、混合CNN-Transformer-Mamba编码器及辅助增强分支。尽管有效,但其是否适合裂缝分割仍存疑问。本文提出将裂缝分割重构为稀疏结构恢复任务:裂缝类别语义有限,但具有强形态规律——细长、稀疏、各向异性、局部断裂,易与纹理或阴影混淆。核心挑战在于保留弱结构信号、恢复方向连贯性、抑制背景干扰。为此,提出RIFT,一套结构对齐的轻量级裂缝分割模型家族。RIFT设计简洁,通过保留局部证据、聚合协同方向连贯性、轻量级多尺度融合实现结构重建。在四个公开基准测试中,RIFT在16项主要指标上达到最佳或并列最佳表现;其中RIFT-B整体精度最强,而RIFT-T仅含0.47M参数,兼具高推理速度与高效部署能力。拓扑感知评估、消融实验、迁移实验与可视化进一步验证:当先验假设贴合裂缝形态时,任务对齐的简化设计可媲美甚至超越复杂混合架构。代码已开源。
原文摘要 · Abstract (English)
Recent crack segmentation methods often follow generic semantic segmentation designs, using stronger backbones, hybrid CNN-Transformer-Mamba encoders, and auxiliary enhancement branches. Although effective, this raises whether stronger generic feature mixing is the most suitable direction for crack segmentation. We instead formulate crack segmentation as sparse structural recovery. Cracks have limited category-level semantics but strong morphological regularities, being thin, sparse, anisotropic, locally fragmented, and easily confused with textures or shadows. Thus, the key bottleneck lies in preserving weak structural evidence, recovering directional continuity, and suppressing background coupling. We propose RIFT, a compact family of morphology-aligned crack segmentation models. Rather than compressing a complex generic architecture, RIFT is simple by design, preserving local evidence, aggregating cooperative directional continuity, and restoring crack structures through lightweight multi-scale fusion. Experiments on four public benchmarks show that RIFT achieves the best or tied-best results across the 16 main metrics against reproduced representative baselines. RIFT-B gives the strongest overall accuracy, while RIFT-T provides the best deployment efficiency with only 0.47M parameters and high inference speed. Topology-aware evaluation, ablations, transfer experiments, and visualizations further verify that task-aligned simplicity can match or surpass complex hybrid architectures when its inductive bias fits crack morphology. Code: https://github.com/xauat-liushipeng/RIFT
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。