从原始网页自动构建高质量指令数据,无需依赖种子数据或结构假设。
Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
- 将网页内容视为指令或响应,通过双视角重构生成指令数据。
- 在4个基准上性能超越现有方法最高16.65%,且数据效率更高。
- 适合需要低成本、高可扩展性数据合成的研究者与工程师。
大语言模型指令遵循能力的提升关键依赖于高质量指令-响应对的可用性。现有自动数据合成方法虽减轻了人工标注负担,但通常严重依赖种子数据质量或对网页结构与内容做出强假设。为此,我们提出完全自动化的Web重建(WebR)框架,直接从原始网页文档中合成高质量指令微调数据,仅需最小假设。利用原始网页内容的内在多样性,我们将网页重建构想为一种新型双视角范式——‘网页作为指令’与‘网页作为响应’,使每个网页被指定为指令或响应以触发重构过程。大量实验表明,WebR生成的数据集在四个指令遵循基准上相比最优基线最高提升16.65%。尤为突出的是,WebR展现出优异的兼容性、数据效率与可扩展性,可低投入实现更强的领域适应。数据与代码已公开于https://github.com/YJiangcm/WebR。
原文摘要 · Abstract (English)
The improvement of LLMs' instruction-following capabilities depends critically on the availability of high-quality instruction-response pairs. While existing automatic data synthetic methods alleviate the burden of manual curation, they often rely heavily on either the quality of seed data or strong assumptions about the structure and content of web documents. To tackle these challenges, we propose Web Reconstruction (WebR), a fully automated framework for synthesizing high-quality instruction-tuning (IT) data directly from raw web documents with minimal assumptions. Leveraging the inherent diversity of raw web content, we conceptualize web reconstruction as an instruction-tuning data synthesis task via a novel dual-perspective paradigm--Web as Instruction and Web as Response--where each web document is designated as either an instruction or a response to trigger the reconstruction process. Comprehensive experiments show that datasets generated by WebR outperform state-of-the-art baselines by up to 16.65% across four instruction-following benchmarks. Notably, WebR demonstrates superior compatibility, data efficiency, and scalability, enabling enhanced domain adaptation with minimal effort. The data and code are publicly available at https://github.com/YJiangcm/WebR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。