对比三种分块策略在土耳其语文档问答中的表现
Comparing Chunking and Embedding Strategies for Turkish RAG Systems

- 采用固定长度、语义和布局感知三种分块方式
- 五种嵌入模型表现无显著差异,快速大模型不更准确
- 布局感知分块对表格密集型文档提升明显
检索增强生成通过从文档集合中检索段落来影响语言模型的输出,其准确性受限于决定可检索内容的分块与嵌入阶段。本文在三个布局迥异的文档上,对比了三种分块策略(固定长度、语义分割、布局感知的Docling)、五种嵌入模型和两种大语言模型在土耳其语文档问答任务上的表现。所有配置均回答相同问题集,通过成对测试分离各组件影响,而非依赖独立基准推断。全交叉实验产生9,000次带评分的问答评估,由独立判别模型评分,组件比较采用配对McNemar检验并经Holm校正。结果表明,三种领先嵌入模型统计上无差异,语言专项化未带来可测量的检索优势;较快的大语言模型并非更准确;最优配置取决于内容类型,布局感知分块对表格密集文档帮助远大于纯文本文档。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation conditions a language model on chunks retrieved from a document collection. Its accuracy is therefore limited by the chunking and embedding stages that determine what can be retrieved. We compare Turkish document question answering across three chunking strategies (fixed-length, semantic, and layout-aware Docling), five embedding models, and two LLMs, over three documents with contrasting layouts. Every configuration answers the same question set, which allows component effects to be separated by paired testing rather than inferred from separate benchmarks. The fully crossed design yields 9{,}000 graded question-answer evaluations, each scored by an independent judge model, and component comparisons are tested by paired McNemar tests under Holm correction. The three leading embedding models are statistically indistinguishable, so language specialization yields no measurable retrieval advantage. The faster LLM is not the more accurate one. The preferred configuration depends on content type, since layout-aware chunking helps table-heavy documents far more than text-heavy ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。