提出评估具身智能数据集多样性和可学习性的新方法。
Data Assessment for Embodied Intelligence
- 构建统一多模态表征,用信息熵衡量数据集信息量。
- 无需训练即可快速评估数据集可学习性,结果准确且可解释。
- 适用于仿真与真实场景数据,帮助优化数据设计。
在具身智能中,数据集起着核心作用,既是知识库也是信息传递的通道。数据集最关键的两个属性是其包含的信息量以及模型学习这些信息的难易程度。然而,具身数据的多模态特性使得评估这些属性尤为困难。以往工作主要关注任务和场景的多样性,通常通过统计任务数量或独立评估单个模态,无法全面反映数据集多样性;而对数据集可学习性的评估则长期被忽视,通常依赖模型训练后进行,过程耗时且缺乏可解释性,难以指导数据改进。本文提出两种基于数据驱动的系统性工具:首先,为每个数据样本构建统一的多模态表示,并在此基础上定义‘多样性熵’,作为连续指标刻画数据集所含信息量;其次,提出首个可解释、数据驱动的算法,在无需训练的前提下高效量化数据集可学习性,使研究者可在数据发布后立即评估其学习潜力。我们在仿真与真实世界具身数据集上验证了该方法的有效性,结果表明其能提供忠实且可操作的洞察,支持同时提升数据集的多样性和可学习性。我们希望这项工作为设计高质量数据集奠定基础,推动具身智能的发展。
原文摘要 · Abstract (English)
In embodied intelligence, datasets play a pivotal role, serving as both a knowledge repository and a conduit for information transfer. The two most critical attributes of a dataset are the amount of information it provides and how easily this information can be learned by models. However, the multimodal nature of embodied data makes evaluating these properties particularly challenging. Prior work has largely focused on diversity, typically counting tasks and scenes or evaluating isolated modalities, which fails to provide a comprehensive picture of dataset diversity. On the other hand, the learnability of datasets has received little attention and is usually assessed post-hoc through model training, an expensive, time-consuming process that also lacks interpretability, offering little guidance on how to improve a dataset. In this work, we address both challenges by introducing two principled, data-driven tools. First, we construct a unified multimodal representation for each data sample and, based on it, propose diversity entropy, a continuous measure that characterizes the amount of information contained in a dataset. Second, we introduce the first interpretable, data-driven algorithm to efficiently quantify dataset learnability without training, enabling researchers to assess a dataset's learnability immediately upon its release. We validate our algorithm on both simulated and real-world embodied datasets, demonstrating that it yields faithful, actionable insights that enable researchers to jointly improve diversity and learnability. We hope this work provides a foundation for designing higher-quality datasets that advance the development of embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。