构建细粒度长文生成数据集,提升文章逻辑与信息深度。
DeFine: A Decomposed and Fine-Grained Annotated Dataset for Long-form Article Generation
- 分层拆解生成流程,引入多级标注与领域知识
- 在三个基线模型上实现主题覆盖与内容忠实度显著提升
- 适合研究长文本生成、可控内容创作的学者使用
长篇文章生成(LFAG)面临逻辑一致性、主题覆盖全面性与叙事连贯性等挑战。现有数据集普遍缺乏层级结构与细粒度标注,导致生成内容浅显且杂乱。为此,我们提出DeFine:一个分层分解且细粒度标注的长篇文章生成数据集。DeFine采用分层分解策略,融合领域知识与多级标注,实现对生成过程的精细控制与内容深度增强。数据构建采用多智能体协作流水线,将生成过程系统化分为四个阶段:数据挖掘、引用检索、问答标注与数据清洗。为验证其有效性,我们设计并测试了三种LFAG基线模型:网络检索、本地检索与基于参考的生成。基于DeFine训练集微调Qwen2-7b-Instruct模型,实验结果显示,在主题覆盖、信息深度和内容忠实度方面均有显著提升。该数据集已公开,以促进后续研究。
原文摘要 · Abstract (English)
Long-form article generation (LFAG) presents challenges such as maintaining logical consistency, comprehensive topic coverage, and narrative coherence across extended articles. Existing datasets often lack both the hierarchical structure and fine-grained annotation needed to effectively decompose tasks, resulting in shallow, disorganized article generation. To address these limitations, we introduce DeFine, a Decomposed and Fine-grained annotated dataset for long-form article generation. DeFine is characterized by its hierarchical decomposition strategy and the integration of domain-specific knowledge with multi-level annotations, ensuring granular control and enhanced depth in article generation. To construct the dataset, a multi-agent collaborative pipeline is proposed, which systematically segments the generation process into four parts: Data Miner, Cite Retreiver, Q&A Annotator and Data Cleaner. To validate the effectiveness of DeFine, we designed and tested three LFAG baselines: the web retrieval, the local retrieval, and the grounded reference. We fine-tuned the Qwen2-7b-Instruct model using the DeFine training dataset. The experimental results showed significant improvements in text quality, specifically in topic coverage, depth of information, and content fidelity. Our dataset publicly available to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。