用数字孪生替代离散分词,让大模型理解真实世界的物理规律。
Position: Foundation Models Need Digital Twin Representations
- 提出用数字孪生作为基础模型的新型表示方式
- 可显式编码领域知识并保持连续动态特性
- 适合需要因果推理与跨模态一致性的场景
当前基础模型依赖直接将连续多模态数据切分为离散标记的表示方式,仅能通过统计相关性学习世界知识,难以维持跨模态语义一致性、捕捉精细时空动态和进行因果推理。这些局限无法通过单纯扩大模型规模或增加数据集解决。本文主张机器学习社区应考虑数字孪生(DT)表示——以结果为导向的数字化身,作为物理过程的虚拟复制品构建块,替代现有分词表示。数字孪生能提供基于物理的表示,显式编码领域知识,保留真实世界过程的连续性,从而有效应对上述挑战。
原文摘要 · Abstract (English)
Current foundation models (FMs) rely on token representations that directly fragment continuous real-world multimodal data into discrete tokens. They limit FMs to learning real-world knowledge and relationships purely through statistical correlation rather than leveraging explicit domain knowledge. Consequently, current FMs struggle with maintaining semantic coherence across modalities, capturing fine-grained spatial-temporal dynamics, and performing causal reasoning. These limitations cannot be overcome by simply scaling up model size or expanding datasets. This position paper argues that the machine learning community should consider digital twin (DT) representations, which are outcome-driven digital representations that serve as building blocks for creating virtual replicas of physical processes, as an alternative to the token representation for building FMs. Finally, we discuss how DT representations can address these challenges by providing physically grounded representations that explicitly encode domain knowledge and preserve the continuous nature of real-world processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。