DART用一张图实现绳索损伤全链条智能评估,无需微调即可输出严重度、建议等多任务结果。
DART: A Vision-Language Foundation Model for Comprehensive Rope Condition Monitoring

- 采用视觉语言融合架构,通过损伤感知掩码和类别自适应门控实现跨模态对齐。
- 在14类损伤上达93.2%准确率,连续严重度预测相关性高达0.94,少样本识别达89.2%。
- 适合海上、工业等场景的智能绳索检测系统开发,可直接部署于现有质检流程。
合成纤维绳索在海上、海洋及工业场景中的状态监测(CM)需要超越分类:需提供连续严重度估计、维护建议、异常标记、劣化时间线及自动化报告,全部基于单张检查图像。本文提出DART(Damage Assessment via Rope Transformer),一种视觉-语言基础模型,通过统一多任务架构解决绳索检测全流程。DART将联合嵌入预测架构(JEPA)扩展至跨模态领域,通过严重度条件交叉模态融合(SC-CMF)模块连接ViT-H/14与Llama-3.2-3B-Instruct。三项创新提升模型泛化能力:(1) HD-MASK,聚焦损伤密集区域的显著性引导掩码;(2) 每类可学习的严重度门控,按损伤类别自适应加权语言对齐;(3) 对比损伤解耦(CDD)损失,使嵌入空间同时编码损伤类型、严重度排序与跨模态语义。在4,270张图像、14种细粒度损伤类别上训练一次后,冻结的DART主干可零微调支持下游任务:损伤分类准确率93.22%,宏平均F1为91.04%(较纯视觉基线提升38.5个百分点),连续严重度回归的斯皮尔曼相关系数达0.94,1阶序数准确率达99.6%;少样本识别在20样本下宏平均F1达89.2%。结果表明,DART可作为通用状态监测骨干,从单一共享表征中提供超越分类的可操作检测智能。
原文摘要 · Abstract (English)
The condition monitoring (CM) of synthetic fibre ropes (SFRs) used in offshore, maritime, and industrial settings demands more than a classifier: inspectors need continuous severity estimates, maintenance recommendations, anomaly flags, deterioration timelines, and automated reports, all from a single inspection image. We present DART (Damage Assessment via Rope Transformer), a vision-language foundation model that addresses the full rope inspection workflow through a unified multi-task architecture. DART extends the Joint-Embedding Predictive Architecture (JEPA) to the cross-modal domain by coupling a Vision Transformer (ViT-H/14) with Llama-3.2-3B-Instruct via a Severity-Conditioned Cross-Modal Fusion (SC-CMF) module. Three architectural innovations drive the model's versatility: (1) HD-MASK, a saliency-guided masking strategy that focuses self-supervised reconstruction on damage-dense patches; (2) per-class learnable severity gates that adaptively weight language grounding by damage category; and (3) a Contrastive Damage Disentanglement (CDD) loss that shapes the embedding space to simultaneously encode damage type, severity ordering, and cross-modal semantics. Trained once on 4,270 images spanning 14 fine-grained rope damage classes, the frozen DART backbone supports downstream tasks without any task-specific fine-tuning: damage classification (93.22 % accuracy, 91.04 % macro-F1, +38.5 pp over a vision-only baseline), continuous severity regression (Spearman rho = 0.94, within-1-ordinal accuracy 99.6 %), few-shot recognition (89.2 % macro-F1 at 20 shots). These results demonstrate that DART functions as a general-purpose CM backbone that goes well beyond classification, providing actionable inspection intelligence from a single shared representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。