验证神经网络是否真正理解并使用物理定律,而非仅靠数据拟合。
LAWFUL: Law-Aligned Witness for Faithful Use of Latents

- 构建因果一致性度量与有效域测试,识别模型内部是否遵循物理规律。
- 在动作捕捉与雷达数据中成功验证多普勒频移定律的内在使用。
- 适合关注模型可解释性与物理规律学习的研究者参考。
当神经网络能准确预测物理系统时,它是否已将支配规律作为形式化、结构化的知识习得?其内部计算是否在整个有效域内都实际运用该表示?我们指出四个阻碍回答这些问题的可解释性缺口:缺乏对连续反事实的覆盖率感知因果一致性度量;缺乏对识别出电路的有效域检验;缺乏对物理规律守恒量与禁止行为的验证;以及缺乏对推导物理量在电路中流动情况的量化。本文提出基础框架 LAWFUL,解决了前两个缺口,并为后两个奠定基础。以 Mocap2Radar 变换器为例,验证了模型是否从动作捕捉和雷达数据中学习并内部使用了多普勒频率定律 $f(t) = \frac{2 v(t)}λ$,其中 $f(t)$ 和 $v(t)$ 均未直接出现。
原文摘要 · Abstract (English)
When a neural network predicts a physical system accurately, has it learned the governing law as formal, structured knowledge, and if so, does the network's internal computation actually use that representation throughout the law's domain of validity? We identify four interpretability gaps that limit answering these questions for {\em physics laws over continuous variables}: the absence of a coverage-aware causal-consistency measure over continuous counterfactuals; of a domain-of-validity test for the identified circuit; of a verification of the law's invariants and forbidden behaviors; and of a quantification of how a derived physical quantity flows through the circuit. We develop a foundational framework, LAWFUL, that closes the first two and lays groundwork for the remaining two, and illustrate it on the Mocap2Radar transformer, validating whether it learns and internally uses the Doppler frequency law $f(t) = \frac{2 v(t)}λ$ from motion-capture and radar data in which neither $f(t)$ nor $v(t)$ appears.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。