arXiv:2605.02884cs.LG2026-05

用无监督学习发现欧洲区域经济结构异常,识别出发展失衡与特殊城市群。

Unsupervised Machine Learning for Detecting Structural Anomalies in European Regional Statistics

论文配图:Unsupervised Machine Learning for Detecting Structural Anomalies in European Regional Statistics
图 1 · 摘自论文原文
  • 构建4个指标的跨区域数据集,融合5种异常检测方法综合判断
  • 识别出10个结构性异常区域,包括发达都市和落后地区
  • 结果反映真实结构差异,可辅助政策分析而非仅查数据错误

确保区域社会经济统计的一致性是各国统计局的核心任务。传统验证工具如范围编辑、比率检查或单变量异常检测虽能识别个别序列的极端值,但在高维情境下难以发现指标组合的异常。本文提出一种无监督机器学习框架,利用公开的欧统局(Eurostat)数据,识别欧洲地区在2022年NUTS2层级上的结构性异常。构建涵盖人均GDP(PPS)、失业率、高等教育完成率及人口密度四个关键指标的横截面数据集,应用并比较五种异常检测方法:单变量z分数、马氏距离、孤立森林、局部离群因子和一类SVM,若一个区域被至少三种方法标记,则视为结构性异常。结果显示,机器学习方法识别出一组与欧盟整体模式显著偏离的区域,包括高度发达的都会经济体(如布鲁塞尔、维也纳、柏林、布拉格),以及长期处于社会经济劣势的地区(如中西部斯洛伐克、匈牙利北部、卡斯蒂利亚-拉曼恰、埃斯特雷马杜拉),还有伊斯坦布尔——其特征明显区别于欧盟首都区域。这些异常并不一定代表数据质量问题,而是反映真实结构差异,值得进一步分析或政策关注。该框架完全可复现、可扩展,且兼容现有验证流程,为欧洲统计体系中早期发现异常区域配置提供灵活工具。

原文摘要 · Abstract (English)

Ensuring the coherence of regional socio-economic statistics is a central task for national statistical institutes. Traditional validation tools, such as range edits, ratio checks, or univariate outlier detection, are effective for identifying extreme values in individual series but are less suited for detecting unusual combinations of indicators in high-dimensional settings. This paper proposes an unsupervised machine learning framework for identifying structurally atypical regional profiles within Europe using publicly available Eurostat data. We construct a cross-sectional dataset of NUTS2 regions (2022) covering four key indicators: GDP per capita in PPS, unemployment rate, tertiary educational attainment, and population density. We apply and compare five anomaly detection techniques, univariate z-scores, Mahalanobis distance, Isolation Forest, Local Outlier Factor, and One-Class SVM, and classify a region as a structural anomaly if it is flagged by at least three of the five methods. The findings show that machine learning methods identify a consistent set of regions whose multivariate profiles diverge substantially from the EU-wide pattern. These include both highly developed metropolitan economies (Brussels, Vienna, Berlin, Prague) and regions with persistent socio-economic disadvantages (Central and Western Slovakia, Northern Hungary, Castilla-La Mancha, Extremadura), as well as Istanbul, whose profile differs markedly from EU capital regions. Importantly, these anomalies do not necessarily signal data quality issues; rather, they reflect meaningful structural divergence that warrants analytical or policy attention. The proposed framework is fully reproducible, scalable, and compatible with existing validation workflows, offering a flexible tool for early detection of unusual regional configurations within the European Statistical System.

异常检测区域经济无监督学习统计验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。