arXiv:2503.10698cs.CLcs.IR2025-03

提出新方法,让文本采样既多样又有序,效果显著优于现有技术。

Ordered Semantically Diverse Sampling for Textual Data

  • 基于嵌入向量主成分生成有序多样化文本样本
  • 在标准分类任务上提升6%至61%的采样效率
  • 适合需要高质量小样本数据的自然语言处理场景

多样性采样的目标是选取一个基数小但信息量大的数据子集。本文提出基于新度量的有序多样化采样问题,设计一种针对文本数据的新型方法,利用嵌入向量的主成分生成有序多样化样本。该方法简单有效,与现有方法相比,在新度量下表现更优。我们将标准文本分类基准转化为有序多样化采样基准。实验表明,当前主流方法性能落后6%至61%,且计算更慢。消融实验验证了各模块对整体指标的贡献。

原文摘要 · Abstract (English)

The goal of diversity sampling is to select a representative subset of data in a way that maximizes information contained in the subset while keeping its cardinality small. We introduce the ordered diverse sampling problem based on a new metric that measures the diversity in an ordered list of samples. We present a novel approach for generating ordered diverse samples for textual data that uses principal components on the embedding vectors. The proposed approach is simple and compared with existing approaches using the new metric. We transform standard text classification benchmarks into benchmarks for ordered diverse sampling. Our empirical evaluation shows that prevailing approaches perform 6% to 61% worse than our method while also being more time inefficient. Ablation studies show how the parts of the new approach contribute to the overall metrics.

文本采样多样性嵌入向量主成分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。