用大模型生成空间分类标签,高效扩展语言场景覆盖
Quantifying and extending the coverage of spatial categorization data sets
- 用大模型生成标签,匹配人类标注结果
- 新增42个场景,覆盖范围优于此前两次扩展
- 适合语言多样性研究与跨语言空间数据构建
不同语言中的空间分类差异常通过人类对拓扑关系图片系列(TRPS)中场景关系的标注来研究。我们证明大语言模型(LLMs)生成的标签与人类标签具有较好一致性,并展示如何利用这些标签决定在现有空间数据集中新增哪些场景和语言。为验证该方法,我们向TRPS增加了42个新场景,结果显示该扩展在可能场景空间的覆盖度上优于以往两次扩展。研究结果为构建包含数十种语言、数百个场景的空间数据集提供了基础。
原文摘要 · Abstract (English)
Variation in spatial categorization across languages is often studied by eliciting human labels for the relations depicted in a set of scenes known as the Topological Relations Picture Series (TRPS). We demonstrate that labels generated by large language models (LLMs) align relatively well with human labels, and show how LLM-generated labels can help to decide which scenes and languages to add to existing spatial data sets. To illustrate our approach we extend the TRPS by adding 42 new scenes, and show that this extension achieves better coverage of the space of possible scenes than two previous extensions of the TRPS. Our results provide a foundation for scaling towards spatial data sets with dozens of languages and hundreds of scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。