arXiv:2608.19492cs.LGcs.RO2026-08

验证多模态感知能否在物理动作中互换使用并保持一致,提升模型对复杂动作的可靠执行能力。

Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

  • 提出响应替换与有序执行验证框架,检验不同传感器信息是否可互换且有效
  • 在19个未见表面测试中,触觉与音频数据响应距离比错误配对近4.5倍,且融合后性能更优
  • 揭示执行器设计关键:全程序输出者比仅共享步骤动态者差38倍,适合关注系统可靠性研究者

世界模型常将紧凑的多模态表征作为感知与物理交互的接口,但现有探测方法无法验证不同传感器是否携带相同可执行语义,或该语义在新动作组合中是否保持。本文引入操作能力层级与分离桥接算子替换验证(DBOSC),检验独立训练的模态编译器能否在训练外证据上互换使用其冻结响应表。在Cluster Haptic数据集上,同一未见表面的触觉与加速度表示在响应空间中的距离比错误配对小4.5倍,且该差异覆盖全部19个保留表面;解封保留响应后发现,每个分支预测物理特性均优于整体响应图。随后在具有互补模态盲区的受控弹塑性系统中测试有序执行。在预设预算下,先决条件拒绝堆叠,因冻结执行器无法推进任何未见程序。在收敛预算下,同一秩三响应图成功执行这些程序(基准NMSE 0.18),融合提升双模态表现,16项注册检查中有14项通过;两项失败源于融合信息矩阵的对角约束表现等同于完整矩阵。通过门控验证表明,通行能力取决于执行器而非响应图本身:输出完整程序的执行器相比仅共享每步动态的实体盲预测器,性能差38倍。一个匹配的非唯一性结果说明,仅靠压缩与融合无法确定未知组合规律。这些结果将属性访问、响应替换、融合闭包与有序执行划分为可独立验证的成果。

原文摘要 · Abstract (English)

World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different sensors carry the same executable meaning or whether that meaning survives a new action composition. We introduce an operational capability hierarchy and the Disjoint-Bridge Operator-Substitution Certificate (DBOSC), which asks whether independently trained modality compilers enter a frozen response chart interchangeably on evidence outside their training panels. On Cluster Haptic, audio and acceleration representations of the same unseen surface are 4.5x closer in response space than wrong-surface pairings, with the gap holding for all 19 held-out surfaces; unsealing withheld responses confirms that every branch predicts the physics better than the population chart. We then test ordered execution in a controlled elastoplastic system with complementary modality blind spots. At the pre-registered budget, the prerequisite refuses the stack because the frozen executor cannot advance even an exact chart coordinate through a held-out program. At a converged budget, the same rank-three chart executes those programs (oracle NMSE 0.18), fusion improves on both modalities, and 14 of 16 registered checks pass; the two failures arise because a diagonal restriction of the fused information matrix performs as well as the full one. Clearing the gate is a property of the executor, not the chart: an executor emitting whole programs instead of shared per-step dynamics is 38x worse than an entity-blind predictor on the same chart. A matching non-identifiability result explains why compression and fusion alone cannot determine an unseen composition law. These results separate attribute access, response substitution, fusion closure, and ordered execution into distinct, separately testable achievements.

多模态物理建模执行验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。