低成本便携设备采集多感官数据,助力AI理解真实世界物体
X-Capture: An Open-Source Portable Device for Multi-Sensory Learning
- 用不到1000美元的设备同步采集视觉、触觉和声音数据
- 收集了500个日常物品的3000组多模态数据点
- 适合做物体感知、跨模态检索等任务的研究者使用
通过多种感官理解物体是人类感知的基础,能实现跨模态融合与更丰富的认知。为使人工智能与机器人系统具备类似能力,获取多样且高质量的多感官数据至关重要。现有数据集往往受限于受控环境、仿真物体或有限的模态组合。我们提出X-Capture,一款开源、便携、低成本(<1000美元)的多感官数据采集设备,可同步捕获相关联的RGBD图像、触觉读数与冲击音频。仅需消费级工具即可组装。利用该设备,我们在多样化真实环境中采集了500个日常物品的3000个数据点,涵盖丰富性和多样性。实验表明,数据的数量与感官广度对物体为中心的任务(如跨模态检索与重建)在预训练与微调中均具显著价值。X-Capture为推进类人多感官表征奠定了基础,强调可扩展性、可及性与现实适用性。
原文摘要 · Abstract (English)
Understanding objects through multiple sensory modalities is fundamental to human perception, enabling cross-sensory integration and richer comprehension. For AI and robotic systems to replicate this ability, access to diverse, high-quality multi-sensory data is critical. Existing datasets are often limited by their focus on controlled environments, simulated objects, or restricted modality pairings. We introduce X-Capture, an open-source, portable, and cost-effective device for real-world multi-sensory data collection, capable of capturing correlated RGBD images, tactile readings, and impact audio. With a build cost under $1,000, X-Capture democratizes the creation of multi-sensory datasets, requiring only consumer-grade tools for assembly. Using X-Capture, we curate a sample dataset of 3,000 total points on 500 everyday objects from diverse, real-world environments, offering both richness and variety. Our experiments demonstrate the value of both the quantity and the sensory breadth of our data for both pretraining and fine-tuning multi-modal representations for object-centric tasks such as cross-sensory retrieval and reconstruction. X-Capture lays the groundwork for advancing human-like sensory representations in AI, emphasizing scalability, accessibility, and real-world applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。