arXiv:2601.03136cs.CLcs.AI2026-01ACL被引 4

揭示视觉语言动作模型数据集语言多样性不足问题

Limited Linguistic Diversity in Embodied AI Datasets

论文配图:Limited Linguistic Diversity in Embodied AI Datasets
图 1 · 摘自论文原文
  • 系统分析多个主流VLA数据集的语言特征
  • 发现指令高度重复,结构变化少,覆盖范围窄
  • 为数据集评估与优化提供可量化的参考基准

语言在视觉-语言-动作(VLA)模型中起关键作用,但训练和评估这些系统的数据集的语言特性仍缺乏系统记录。本文对多个广泛使用的VLA语料库进行系统性审计,旨在刻画数据集中实际包含的指令类型及其语言多样性。我们从词汇多样性、重复与重叠、语义相似性、句法复杂度等维度量化指令语言。分析表明,许多数据集依赖高度重复、模板化的指令,导致指令形式分布狭窄。这些发现旨在作为当前VLA训练与评估数据语言信号的描述性记录,支持更细致的数据集报告、更合理的数据集选择,以及针对性的语料扩充或优化策略。

原文摘要 · Abstract (English)

Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate these systems remain poorly documented. In this work, we present a systematic dataset audit of several widely used VLA corpora, aiming to characterize what kinds of instructions these datasets actually contain and how much linguistic variety they provide. We quantify instruction language along complementary dimensions--including lexical variety, duplication and overlap, semantic similarity, and syntactic complexity. Our analysis shows that many datasets rely on highly repetitive, template-like commands with limited structural variation, yielding a narrow distribution of instruction forms. We position these findings as descriptive documentation of the language signal available in current VLA training and evaluation data, intended to support more detailed dataset reporting, more principled dataset selection, and targeted curation or augmentation strategies that broaden language coverage.

VLA语言多样性数据集审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。