用零售场景视频训练模型,发现外视角数据比内外视角混合效果更好。
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail

- 用同步的内外视角视频数据,对比三种训练方式的效果差异。
- 仅用1.6万条外视角视频,性能超过3.2万条混合数据的模型。
- 外视角数据对模型改进更显著,尤其在短期预测阶段最有效。
基础视频扩散模型被视为具身智能体的世界模拟器,但其在互联网通用视频上预训练,与真实部署场景匹配度低。本文研究如何高效适配预训练视频世界模型至零售场景:当同一活动的同步内视角与外视角视频可用时,哪种视角的数据能生成最强适应模型?我们构建了RetailSMV(零售同步多视角)数据集,包含5家超市中店员视角(补货、陈列、称重、管理推车、结账)的32,105段带标注视频,且同步采集内外视角,而非以往以顾客为中心的零售视频。在相同超参数下,训练了三种匹配的低秩适配(LoRA)配置:仅内视角、仅外视角、内外结合。在200段独立测试集上,采用七项互补指标和严格配对统计协议评估,结果表明:仅使用外视角数据的适配模型在六项指标上达到或超过混合数据模型,并在LPIPS、PSNR和DreamSim上显著更优。进一步对称对比显示,将外视角加入内视角训练有帮助,反之则有害。绝对适应差距在最短轨迹预测时间点最大,表明近域预测是适配收益最高的阶段。
原文摘要 · Abstract (English)
Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces the strongest adapted model? We introduce RetailSMV (Retail Synchronized Multi-View), a corpus of 32,105 captioned retail clips from five supermarkets with synchronized ego/exo capture from the store-staff perspective (stocking, arranging, weighing, managing supply carts, scanning at checkout), rather than the customer-centric framing of prior retail video corpora, and train three matched Low-Rank Adaptation (LoRA) configurations of Cosmos3-Nano (egocentric-only, exocentric-only, combined) under identical hyperparameters. On a 200-clip held-out test set evaluated with seven complementary metrics under a strict paired statistical protocol, exocentric-only adaptation matches or exceeds combined adaptation on six of seven point estimates and is significantly better on LPIPS, PSNR, and DreamSim, despite training on only 15,985 exocentric clips (versus 32,105 for combined). A symmetric paired comparison further shows that adding exocentric data to egocentric-only training helps while adding egocentric data to exocentric-only training hurts. The absolute adaptation gap is largest at the shortest rollout time, identifying the near-horizon prediction window as the regime in which adaptation is most beneficial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。