arXiv:2606.00881cs.CL2026-06被引 1

系统评估多种文本分块方法,揭示其在RAG中的性能与计算成本的权衡。

Chunking Methods on Retrieval-Augmented Generation - Effectiveness Evaluation Against Computational Cost and Limitations

  • 对比多种分块策略,涵盖固定大小与语义分块等方法
  • 发现多数新方法仅在特定场景有效,泛化能力有限
  • 强调分块作为关键预处理步骤的潜在问题与挑战

检索增强生成(RAG)显著提升了大语言模型(LLMs)的性能。其中,文本分块是核心任务之一。传统上,固定大小分块和语义分块是主流方法。然而,随着对分块策略兴趣上升,大量新方法被提出,常宣称优于传统方法。这些方法多针对特定场景与数据类型,缺乏在多样化场景下的有效性证据。因此,不同方法间的直接比较仍具挑战性。据我们所知,本研究首次系统评估了广泛分块方法的有效性,并揭示了RAG中分块策略的底层难题。尽管分块常被视为简单预处理步骤,我们表明其引入了一系列重要且常被忽视的问题。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) has demonstrated significant capabilities in enhancing the performance of Large Language Models (LLMs). One of the key tasks in RAG systems is the chunking process. Traditionally, fixed-size chunking and semantic chunking have been the standard approaches. However, interest in chunking strategies has been increasing, leading to a growing number of proposed methods that often claim improved performance over these conventional techniques. Many of these approaches are tailored to specific use cases and data types, with limited evidence of their effectiveness across diverse scenarios. As a result, it remains challenging to directly compare different techniques and assess their relative strengths. To the best of our knowledge, this study is the first to systematically evaluate the effectiveness of a wide range of chunking methods and emphasize the underlying challenges of chunking strategies in RAG systems. While chunking is commonly treated as a simple preprocessing step, we show that it introduces a range of impactful and often overlooked issues.

RAG文本分块效能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。