用测试环境的多模态数据自监督训练,让模型更适应特定场景。
Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality

- 在测试环境中收集多模态数据,通过跨模态预测自监督预训练。
- 仅用测试环境数据训练的模型,性能媲美互联网大规模数据训练的通用模型。
- 适合部署于传感器丰富但无外部数据资源的专用设备,如家庭机器人。
跨模态学习(即从一种模态预测另一种)是利用多模态实现自监督学习的基本机制。许多实际应用(如家用机器人部署)中的设备具备丰富传感器,可在测试环境中实现多模态感知。这为利用设备在测试环境中采集的多模态数据进行表示学习提供了机会。发展心理学研究也表明生物体利用此机制构建对环境的有效表征。为此,我们设计了一个受控实验设置:将用户设备限制在特定测试环境中,形成一种专属性训练设定。在此框架下,提出测试空间训练(Test-Space Training, TST),在测试环境中收集多模态数据并进行自监督预训练。在相同环境下评估下游任务表现,发现仅使用测试环境的丰富多模态数据,并结合跨模态学习,即可达到与在大规模互联网数据上预训练的通用模型(如 DINOv2、CLIP)相媲美的性能。这提供了一种减少对互联网级数据依赖的替代方案。此外,通过分析与消融实验揭示了以多模态替代数据的潜力,以及不同预训练数据如何在模型专属性与泛化性之间形成权衡。
原文摘要 · Abstract (English)
Cross-modal learning, i.e., learning to predict one modality from another, is a fundamental mechanism for self-supervision via leveraging multimodality. Many practical applications, e.g., deploying a household robot, involve devices that are equipped with a rich set of sensors that enable multimodal sensing in their test environment. This presents an opportunity to apply cross-modal learning to the multimodal data sensed by these devices to learn representations. Findings in developmental psychology also suggest that biological agents leverage it to build an effective representation of their surroundings. To study this, we propose a controlled setup, where we restrict a user device to just a given test environment. It results in a specialization setup where we attempt to develop a performant model for this specific test environment. Under this setup, we develop Test-Space Training (TST), which performs multimodal data collection in the test environment and performs self-supervised pre-training on it. We evaluate these models on various downstream tasks in the same environment. Under this setup, we find various interesting insights, such as collecting rich multimodal data only from the test environment and leveraging cross-modal learning, we can achieve competitive results with generalist models (e.g., DINOv2 and CLIP) pre-trained on large-scale internet datasets. This enables an alternative scenario where the need for external Internet-scale datasets for pre-training models is reduced. We also present a set of analyses and ablations that raise intriguing points on substituting data with (multi)modality, and how varying pre-training data enables a tradeoff between a model's abilities to specialise to a test environment, and generalize to held-out spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。