发现科学模型因离散分词导致几何失真,提出连续头可显著改善
The Geometric Alignment Tax: Tokenization vs. Continuous Geometry in Scientific Foundation Models
- 用连续输出头替代离散分词,减少8.5倍几何畸变
- 细粒度量化反而加剧几何失真,存在非单调陷阱
- 适合研究生物物理建模中表示几何保真的学者
用于生物和物理领域的基础模型虽优化了预测准确率,但其内部表征系统性地破坏了所建模系统的连续几何结构。我们识别出根本原因:几何对齐税——将连续流形强行通过离散分类瓶颈所导致的内在代价。在合成动力系统上的受控消融实验表明,将交叉熵替换为连续输出头,在相同编码器下可使几何失真降低最多8.5倍;而学习的码本表现出非单调双重困境:更细粒度的量化虽提升重建质量,却恶化几何结构。在连续目标下,三种架构差异仅1.3倍;而在离散分词下,差异达3000倍。通过率失真理论与MINE评估14个生物基础模型,识别出三种失效模式:局部-全局解耦、表征压缩、几何空洞。受控实验证实,Evo 2在真实DNA上的逆互补鲁棒性反映的是保守序列组成,而非学习到的对称性。无一模型能同时实现低失真、高互信息与全局一致性。
原文摘要 · Abstract (English)
Foundation models for biology and physics optimize predictive accuracy, but their internal representations systematically fail to preserve the continuous geometry of the systems they model. We identify the root cause: the Geometric Alignment Tax, an intrinsic cost of forcing continuous manifolds through discrete categorical bottlenecks. Controlled ablations on synthetic dynamical systems demonstrate that replacing cross-entropy with a continuous head on an identical encoder reduces geometric distortion by up to 8.5x, while learned codebooks exhibit a non-monotonic double bind where finer quantization worsens geometry despite improving reconstruction. Under continuous objectives, three architectures differ by 1.3x; under discrete tokenization, they diverge by 3,000x. Evaluating 14 biological foundation models with rate-distortion theory and MINE, we identify three failure regimes: Local-Global Decoupling, Representational Compression, and Geometric Vacuity. A controlled experiment confirms that Evo 2's reverse-complement robustness on real DNA reflects conserved sequence composition, not learned symmetry. No model achieves simultaneously low distortion, high mutual information, and global coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。