arXiv:2607.19848cs.CL2026-07被引 1

提供基于嵌入的标准化数据多样性测量工具,助力公平稳健的NLP模型开发。

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

论文配图:emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity
图 1 · 摘自论文原文
  • 通过嵌入向量量化文本的风格、语义、语言和说话人多样性
  • 支持任意嵌入模型与可嵌入数据,覆盖多种多样性维度
  • 开源工具包适用于数据质量评估与模型鲁棒性研究

现有研究表明,数据多样性对构建公平且稳健的自然语言处理模型至关重要。然而当前多样性度量方法不一致且碎片化:虽然已有多种文本词汇多样性测量工具,但缺乏基于嵌入的标准化度量手段。嵌入式多样性度量具有高度灵活性,可适配任意嵌入模型和可嵌入数据,适用于多种多样性概念。本文提出 emb-diversity 工具,涵盖广泛多样性的测量方法。我们展示了其在多个应用场景中的潜力:衡量数据集的风格、语义、语言及说话人多样性。项目开源地址:https://github.com/nlpsoc/emb-diversity/

原文摘要 · Abstract (English)

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmented: While there exist a number of tools for measuring the lexical diversity of texts, researchers lack standardized tools for quantifying diversity based on embeddings. Embedding-based diversity measures are highly flexible: They work with any embedding model and any data that can be embedded, and are thus applicable to many notions of diversity. With emb-diversity, we provide a comprehensive embedding-based diversity measurement tool, spanning a broad range of measures. We demonstrate its potential for several use cases: measuring the stylistic, semantic, language and speaker diversity of datasets. https://github.com/nlpsoc/emb-diversity/

数据多样性嵌入度量NLP评估开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。