arXiv:2409.17827cs.CL2024-09NeurIPS被引 4

构建1590亿词的商业文本数据集,更真实且更少毒性。

BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text

  • 从企业披露文件中提取1590亿词,构建大规模商业语料
  • 相比主流网络数据,毒性降低18%-33%,金融领域表现更优
  • 适合训练低毒性、高专业性的大模型,尤其金融领域

近期语言建模的突破主要源于在更大规模数据集上扩展相同模型架构。本文提出BeanCounter,一个公开的商业文本数据集,包含超过1590亿个标记(tokens),数据源自企业披露文件。该数据与基于Common Crawl的数据重合度低于0.1%,规模比同类来源的数据大一个数量级。由于其来源特性,我们假设BeanCounter相较网络数据更真实、毒性更低。实验发现,多个社会身份在其中出现频率相近,但上下文毒性显著低于其他数据集。通过持续预训练两个LLM,我们观察到毒性强生成减少18%-33%,且在金融任务上的性能提升明显。结果表明,BeanCounter是可支持数十亿参数模型训练的高质量、低毒性、领域专一的数据源。

原文摘要 · Abstract (English)

Many of the recent breakthroughs in language modeling have resulted from scaling effectively the same model architecture to larger datasets. In this vein, recent work has highlighted performance gains from increasing training dataset size and quality, suggesting a need for novel sources of large-scale datasets. In this work, we introduce BeanCounter, a public dataset consisting of more than 159B tokens extracted from businesses' disclosures. We show that this data is indeed novel: less than 0.1% of BeanCounter appears in Common Crawl-based datasets and it is an order of magnitude larger than datasets relying on similar sources. Given the data's provenance, we hypothesize that BeanCounter is comparatively more factual and less toxic than web-based datasets. Exploring this hypothesis, we find that many demographic identities occur with similar prevalence in BeanCounter but with significantly less toxic context relative to other datasets. To demonstrate the utility of BeanCounter, we evaluate and compare two LLMs continually pre-trained on BeanCounter with their base models. We find an 18-33% reduction in toxic generation and improved performance within the finance domain for the continually pretrained models. Collectively, our work suggests that BeanCounter is a novel source of low-toxicity and high-quality domain-specific data with sufficient scale to train multi-billion parameter LLMs.

数据集低毒性金融NLP大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。