arXiv:2501.08003cs.CL2025-01

提升法语数据集采样的多样性,用低成本词法多样性替代高成本句法多样性。

Formalising lexical and syntactic diversity for data sampling in French

  • 设计启发式算法,比随机采样显著提升数据多样性。
  • 发现词法与句法多样性相关性随数据集和度量方式变化。
  • 适合关注数据质量与多样性评估的研究者参考。

多样性是数据集的重要属性,基于多样性的采样对数据集构建很有帮助。由于寻找最优多样性样本代价高昂,本文提出一种启发式方法,显著提升了相对于随机采样的多样性。同时探讨了词法多样性与句法多样性之间的相关性,旨在通过低成本的词法多样性来实现高成本的句法多样性采样。研究发现,这种相关性在不同数据集和多样性度量版本间波动明显,表明任意选择的度量可能无法准确捕捉数据集的多样性特征。

原文摘要 · Abstract (English)

Diversity is an important property of datasets and sampling data for diversity is useful in dataset creation. Finding the optimally diverse sample is expensive, we therefore present a heuristic significantly increasing diversity relative to random sampling. We also explore whether different kinds of diversity -- lexical and syntactic -- correlate, with the purpose of sampling for expensive syntactic diversity through inexpensive lexical diversity. We find that correlations fluctuate with different datasets and versions of diversity measures. This shows that an arbitrarily chosen measure may fall short of capturing diversity-related properties of datasets.

数据采样多样性法语词法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。