arXiv:2604.26366stat.MLcs.LG2026-04

用扩散模型评估结构监测数据质量,抗异常值能力强。

Probabilistic data quality assessment for structural monitoring data via outlier-resistant conditional diffusion model

  • 基于条件扩散模型,融合时序上下文与分位数归一化。
  • 每个数据点生成异常概率,整体数据集得分更准确。
  • 适合真实结构监测场景,对异常值鲁棒性强。

数据质量评估是确保后续结构健康监测(SHM)任务可靠性的关键步骤。本文提出一种基于预测偏差的SHM数据质量评估方法,采用单变量隐式自回归模型,实现异常值诊断与数据清洗。所提出的条件扩散模型(CDM)在标准扩散模型基础上引入条件嵌入模块以融入时序上下文,采用四分位数归一化缓解分布偏斜,并使用Huber损失增强对异常值的鲁棒性。在此单变量隐式自回归框架中,每个数据点被赋予异常概率,量化其“异常程度”,并计算全局质量评估分数以表征整个数据集的质量。利用实际结构的运行数据进行的大量案例研究显示,该框架显著提升了数据质量评估的准确性,优于代表聚类、孤立性检测及深度重建方法的多种强基线。消融实验与超参数分析进一步验证了该框架的有效性与鲁棒性。

原文摘要 · Abstract (English)

Data quality assessment is an essential step that ensures the reliability of the subsequent structural health monitoring (SHM) tasks. This study proposes a prediction deviation-based SHM data quality assessment method using a univariate implicit auto-regressive model, enabling outlier diagnosis and data cleaning. The proposed conditional diffusion model (CDM) augments the standard diffusion model with a conditional embedding module to incorporate temporal context, quartile normalization to mitigate distribution skew, and a Huber loss to enhance robustness against outliers. Within this univariate implicit autoregressive framework, each data point is assigned an outlier probability, quantifying its degree of "outlier-ness", and a global quality evaluation score is computed to characterize the overall dataset quality. Extensive case studies utilizing operational data from real-world structures demonstrate that the proposed framework significantly improves the accuracy of data quality assessment, outperforming other strong baselines representative of clustering, isolation-based, and deep reconstruction methods. The effectiveness and robustness of the proposed framework are further demonstrated by the findings of ablation experiments and hyperparameter analysis.

数据质量结构监测扩散模型异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。