arXiv:2511.19647cs.ROcs.AI2025-11

用机器人持续采集真实场景数据,自动优化视觉语言模型。

Robot-Powered Data Flywheels: Deploying Robots in the Wild for Continual Data Collection and Foundation Model Adaptation

  • 机器人边执行任务边收集真实世界数据,形成数据迭代闭环。
  • 在2103个书架上部署后,识别准确率从32%提升至71.8%。
  • 适合需要持续适应现实环境的机器人与大模型研究者。

基础模型(FM)在视觉和语言任务中展现出强大的零样本能力,但其依赖互联网预训练数据,在非结构化真实环境中表现脆弱。部署过程中遇到的杂乱数据(如遮挡或多语言文本)在现有语料库中严重缺失。机器人作为具身智能体,能主动在物理环境中采集大规模真实数据,填补当前模型所缺的典型实例。本文提出机器人驱动的数据飞轮框架,使机器人从基础模型使用者转变为数据生成者。通过在真实场景部署配备基础模型的机器人,实现良性循环:机器人完成实用任务的同时,收集用于模型微调的真实数据,提升特定领域适应性和邻近领域的泛化能力。我们以扫描仪机器人Scanford为例,在东亚图书馆连续运行两周,自主扫描书架、使用视觉-语言模型(VLM)识别图书,并利用馆藏目录自动标注图像,无需人工标注。该部署既协助了图书馆员,又生成可用于微调底层VLM的数据集,显著提升了模型在真实图书馆场景中的性能及跨领域的多语言光学字符识别(OCR)表现。基于2103个书架的数据,模型在图书识别准确率从32.0%提升至71.8%,英语多语言OCR从24.8%升至46.6%,中文从30.8%升至38.0%,同时节省约18.7小时人力。结果表明,机器人驱动的数据飞轮既能降低实际部署中的人力负担,也为基础模型持续适应现实复杂性提供了新路径。

原文摘要 · Abstract (English)

Foundation models (FM) have unlocked powerful zero-shot capabilities in vision and language, yet their reliance on internet pretraining data leaves them brittle in unstructured, real-world settings. The messy, real-world data encountered during deployment (e.g. occluded or multilingual text) remains massively underrepresented in existing corpora. Robots, as embodied agents, are uniquely positioned to close this gap: they can act in physical environments to collect large-scale, real-world data that enriches FM training with precisely the examples current models lack. We introduce the Robot-Powered Data Flywheel, a framework that transforms robots from FM consumers into data generators. By deploying robots equipped with FMs in the wild, we enable a virtuous cycle: robots perform useful tasks while collecting real-world data that improves both domain-specific adaptation and domain-adjacent generalization. We instantiate this framework with Scanford, a mobile manipulator deployed in the East Asia Library for 2 weeks. Scanford autonomously scans shelves, identifies books using a vision-language model (VLM), and leverages the library catalog to label images without human annotation. This deployment both aids librarians and produces a dataset to finetune the underlying VLM, improving performance on the domain-specific in-the-wild library setting and on domain-adjacent multilingual OCR benchmarks. Using data collected from 2103 shelves, Scanford improves VLM performance on book identification from 32.0% to 71.8% and boosts domain-adjacent multilingual OCR from 24.8% to 46.6% (English) and 30.8% to 38.0% (Chinese), while saving an ~18.7 hrs of human time. These results highlight how robot-powered data flywheels can both reduce human effort in real deployments and unlock new pathways for continually adapting FMs to the messiness of reality. More details are at: https://scanford-robot.github.io

机器人数据飞轮视觉语言模型持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。