arXiv:2605.09677cs.CV2026-05被引 1

无需训练、标记和标定,用视觉大模型实现结构位移精准测量

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement

论文配图:VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement
图 1 · 摘自论文原文
  • 基于视觉大模型推断相机参数并跟踪点,通过三角测量重建位移
  • 垂直/横向位移误差低(NRMSE 0.11/0.12),相关性达0.86/0.88
  • 适合数字孪生与智能建造中的自动化结构监测场景

可靠的位移测量是结构健康监测与数字工程流程的基础,提供直接的结构响应信息。基于视觉的测量方法因其低成本、非接触特性成为有前景的方案,但部署常受限于任务特定模型训练或现场准备,如标记安装或手动相机标定。本研究提出一种基于视觉基础模型的结构位移测量框架(VFM-SDM),融合VFM推断的相机参数估计与点跟踪,通过三角测量重建多方向位移,无需任务特定训练或现场准备,支持真实场景下的高效非接触部署。引入结构几何约束以抑制物理上不合理偏差,提升估计一致性。构建了来自在役人行桥的多模态数据集,并提出统一基准评估协议以支持可复现评估。代表性结果表明,垂直与横向位移的低幅值误差(NRMSE$_{\text{range}}$: 0.11/0.12)、强时间一致性(相关系数: 0.86/0.88)及小峰值幅值误差(RPPAE: 0.01/0.02),证明其在真实条件下的鲁棒性能。该框架推动了自动化、可扩展的位移监测发展,为视觉大模型赋能的结构响应测量在数字孪生与数据驱动建设流程中奠定基础。

原文摘要 · Abstract (English)

Reliable displacement measurement is fundamental for structural health monitoring and digital engineering workflows, as it provides direct structural response information. Vision-based measurement has emerged as a promising approach for low-cost, non-contact displacement monitoring. However, its deployment often remains constrained by task-specific model training or on-site preparation, such as marker installation or manual camera calibration. This study presents a Vision Foundation Model-based framework for Structural Displacement Measurement (VFM-SDM) that integrates VFM-inferred camera parameter estimation and point tracking to reconstruct multi-directional structural displacements via triangulation without task-specific training or on-site preparation, enabling efficient non-contact deployment in real-world applications. Structural geometry constraints are incorporated to suppress physically implausible deviations and improve estimation consistency. A multi-modal field dataset collected from an in-service pedestrian bridge is introduced alongside a unified benchmarking protocol to support reproducible evaluation. Representative results show low amplitude errors (NRMSE$_{\text{range}}$: 0.11/0.12), strong temporal agreement (correlation coefficient: 0.86/0.88), and small peak-to-peak amplitude errors (RPPAE: 0.01/0.02) for vertical and lateral displacements, indicating robust performance under real-world conditions. The proposed framework advances automated, scalable displacement monitoring and lays the groundwork for VFM-enabled structural response measurements in digital twin and data-centric construction workflows.

结构监测视觉大模型非接触测量数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。