arXiv:2607.04563cs.CL2026-07

提出新指标,量化文本数据的忠实度与多样性,助力高质量数据筛选。

Fidelity-Diversity Metrics for Text

论文配图:Fidelity-Diversity Metrics for Text
图 1 · 摘自论文原文
  • 基于最优传输理论,构建衡量文本忠实度与多样性的双指标。
  • 在M2D2数据集上区分出文本缺失忠实度或多样性的不同问题。
  • 发现合成数学题数据多样性不足会降低微调模型的准确率。

随着语言建模技术的发展,研究者愈发关注训练数据集的构建与筛选。例如,从业者常通过扩充高质量数据集来提升模型性能。但这一过程需要更精细的数据质量评估。本文借鉴生成模型精度与召回率的度量思路,提出一对新指标:(1) 忠实度,衡量候选文本与参考数据的相似程度;(2) 多样性,衡量其对参考数据分布模式的覆盖程度。该指标基于离散文本摘要间的最优传输分歧函数构建。在M2D2文本数据集上的实验表明,该方法能有效区分候选文本在忠实度与多样性上的缺陷。进一步实验发现,合成的GSM8K风格数学数据存在多样性不足问题,且该缺陷与微调后语言模型的下游准确率下降显著相关。

原文摘要 · Abstract (English)

As language modeling technology matures, there is an increasing research focus on the composition and curation of datasets used to train these models. For instance, practitioners commonly seek to augment high-quality datasets with additional text to enhance the performance of models trained on that data. However, informed decisions about data augmentation require more nuanced assessments about data quality. We build on work measuring the precision and recall of generative models to develop a pair of metrics that quantify (1) fidelity, capturing how closely candidate text resembles reference data, and (2) diversity, capturing how well it covers the modes of the reference dataset. Our metrics are based on optimal transport divergence functionals between discrete text summaries. In experiments on M2D2 text datasets, we show that these metrics are able to disentangle a lack of fidelity from a lack of diversity in deficient candidate text. In further experiments, our metrics detect diversity deficits in synthetic GSM8K-style math datasets, which correlate with degradations in downstream accuracy of language models finetuned on this synthetic data.

文本生成数据评估最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。