arXiv:2508.01309cs.CLcs.AI2025-08

用大模型自动生成带思维链的高质量问答数据,省时省钱。

D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation

  • 以文档为中心分段处理,结合思维链推理生成问答对
  • 每小时每卡生成超1100条高质量数据,效率极高
  • 适合需要快速构建领域专用数据集的研究者

高质量领域特定问答(QA)数据集稀缺且成本高昂,限制了大语言模型(LLMs)的监督微调。我们提出D-SCoRE,一种无需训练的框架,利用大语言模型和提示工程,从任意文本源自动构建多样、丰富的带思维链(CoT)问答数据集。通过整合文档中心处理、分段、思维链推理与结构化导出,并引入语义角色变换、问题类型平衡、反事实增强等多维控制策略,D-SCoRE生成的问答对在多样性与相关性上均有提升。在多数评估领域中,使用D-SCoRE生成数据微调的LLM性能优于使用人工标注数据训练的模型。其高效可扩展性使得在消费级硬件上实现快速高精度领域适配微调成为可能,端到端每GPU小时生成超过1,100条高质量问答对。

原文摘要 · Abstract (English)

The scarcity and high cost of high-quality domain-specific question-answering (QA) datasets limit supervised fine-tuning of large language models (LLMs). We introduce $\textbf{D-SCoRE}$, a training-free framework that leverages LLMs and prompt engineering to automatically generate diverse, rich QA datasets with Chain-of-Thought (CoT) from arbitrary textual sources. By integrating $\textbf{D}$ocument-centric processing, $\textbf{S}$egmentation, $\textbf{Co}$T $\textbf{R}$easoning, and structured $\textbf{E}$xport - along with multi-dimensional controls such as semantic role transformation, question type balancing, and counterfactual augmentation - D-SCoRE produces tailored QA pairs with enhanced diversity and relevance. LLMs fine-tuned on D-SCoRE-generated datasets outperform those trained on human-annotated QA data across most evaluated domains. Its efficiency and scalability enable rapid, high-performance domain-adaptive fine-tuning on consumer-grade hardware, generating over 1,100 high-quality QA pairs per GPU-hour end-to-end.

数据生成思维链领域适应自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。