提出可控框架OVTF,研究触觉与视觉未来信息如何有效指导动作决策。
Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models

- 构建虚拟成功轨迹的视觉-触觉未来配对数据,隔离接口问题
- 提出的AFM模型使触觉记忆同步感知对应视觉未来,成功率提升至32.0%
- 适合关注多模态预测与动作规划融合的研究者
接触丰富的操作任务仍具挑战性,因成功控制依赖于常难以仅从视觉中捕捉的物理交互线索。现有触觉世界动作模型联合建模未来视觉观测与触觉信号以指导动作生成,但其未来结构如何适配动作专家尚不明确。直接使用学习模型研究该问题困难,因端到端行为会混杂无效视觉未来、不可靠预测、不准确或跨模态不一致的触觉预报,以及难以理解的未来到动作接口。为使该接口可独立研究,我们引入奥义视觉-触觉预知(OVTF),一个在仿真中提供经验证的成功轨迹的图像与触觉未来配对数据的受控框架。通过固定未来生成器,OVTF隔离接口,并提出更清晰的问题:若未来成功且物理可执行,何种表示能让动作专家有效利用?在OVTF中,我们提出非对称相位局部未来记忆(AFM),其中视觉记忆读取未来视觉,每个触觉记忆同时关注自身触觉流和相位对齐的未来视觉,禁止跨触觉访问。与去除视觉到触觉访问的模态隔离未来记忆(IFM)对比,在UniVTAC仿真基准的七项任务上,AFM平均成功率达32.0%,优于IFM的23.7%和UniVTAC-ACT的14.9%。该对照表明,选择性相位对齐的视觉-触觉路由比完全模态隔离更利于构建可行动的未来到动作桥梁。
原文摘要 · Abstract (English)
Contact-rich manipulation remains challenging because successful control depends on physical interaction cues that are often weakly observable from vision alone. Recent tactile world action models jointly model future visual observations and tactile signals to guide action generation, but how such futures should be structured for effective use by the action expert remains underexplored. Directly studying this question with learned world action models is difficult because end-to-end behavior entangles physically invalid visual futures, unreliable predictions, inaccurate or cross-modally inconsistent tactile forecasts, and an unreadable future-to-action interface. To make this interface independently studyable, we introduce Oracle Visuo-Tactile Foresight (OVTF), a controlled framework that supplies paired RGB and tactile futures from successful trajectories verified in simulation. By fixing the future provider, OVTF isolates the interface and asks a cleaner question: if the future is successful and physically executable, what representation allows the action expert to absorb its benefit? Within OVTF, we propose Asymmetric Phase-Local Future Memory (AFM), in which visual memory reads future vision, each tactile memory jointly attends to its own tactile stream and phase-aligned future vision, and cross-tactile access is blocked. We compare AFM with Modality-Isolated Future Memory (IFM), which removes visual-to-tactile access and processes each future modality independently. Across seven tasks on the UniVTAC simulation benchmark, AFM achieves 32.0% average success, compared with 23.7% for IFM and 14.9% for UniVTAC-ACT. This controlled comparison shows that selective phase-aligned visual-tactile routing provides a more actionable future-to-action bridge than complete modality isolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。