arXiv:2606.16737math-phcs.LG2026-06

从数据中自动发现物理系统的无量纲量,无需先验知识。

The Algebra of Units: From Buckingham's Pi-grec Theorem to Latent-Variable Learning

论文配图:The Algebra of Units: From Buckingham's Pi-grec Theorem to Latent-Variable Learning
图 1 · 摘自论文原文
  • 对物理量取对数后,数据落在由无量纲量决定的低维流形上。
  • 用SVD识别流形,再通过整数指数组合恢复出正确无量纲量。
  • 可复现压缩机性能图,误差低于0.01%,适合工程建模与物理学习。

工程师常测量速度、压力、温度、长度等不同单位的物理量。伯金汉姆π定理指出,这些变量总能组合成一组无量纲数,其值完全决定系统行为。传统上,确定合适的无量纲组需依赖专家知识和物理洞察。本文表明,这些无量纲量可仅从数据中自动发现,无需预先了解物理规律。关键观察是:对测量值取对数后,同一系统在不同尺度下的数据点位于一个由底层无量纲量决定的低维流形上。奇异值分解(SVD)可直接从数据中识别该流形。随后通过整数指数组合搜索候选无量纲量,并利用重复变量过滤器仅保留由机器特征尺度构成的量。该方法成功恢复了流量系数、扬程系数、马赫数等经典工程量,同时排除了等价但不易解释的替代项。在包含16,000个测量值的合成压缩机数据集上验证,仅基于原始带量纲变量和无物理输入,即可以数值精度恢复正确的无量纲组,并重建压缩机性能图,误差低于0.01%。更广泛地,该工作揭示了经典量纲分析与现代数据驱动学习之间存在紧密联系,二者均依赖相同的代数结构,为构建兼具可解释性、可扩展性和数据效率的物理模型提供了新思路。

原文摘要 · Abstract (English)

Engineers often measure many quantities-speed, pressure, temperature, length-expressed in different physical units. The Buckingham Pi-grec theorem states that these variables can always be combined into a smaller set of dimensionless numbers whose values fully determine the system's behaviour. Identifying the appropriate dimensionless groups has traditionally required expert knowledge and physical insight. This paper shows that they can instead be discovered automatically from data, without prior knowledge of the governing physics. The key observation is that, after logarithmic transformation, measurements collected under different scalings of the same system lie on a low-dimensional manifold whose geometry is determined by the underlying dimensionless groups. Singular value decomposition (SVD) identifies this manifold directly from data. A subsequent search over integer-exponent combinations recovers candidate dimensionless quantities, while a repeating-variable filter retains only those constructed from the machine's characteristic scales. This procedure recovers familiar engineering groups, including the flow coefficient, head coefficient, and Mach number, while excluding equivalent but less interpretable alternatives. The method is demonstrated on a synthetic compressor dataset containing 16,000 measurements. Starting from raw dimensional variables and no physics input, it recovers the correct dimensionless groups to numerical precision and reproduces the compressor performance map with an error below 0.01%. More broadly, the work reveals a close connection between classical dimensional analysis and modern data-driven learning. Both rely on the same underlying algebraic structure, suggesting new approaches for building physical models that are simultaneously interpretable, scalable, and data-efficient.

量纲分析数据驱动无量纲量SVD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。