用大模型自动构建小众领域的高质量内容库,提升数据采集效率。
Leveraging LLMs to Create Content Corpora for Niche Domains
- 用大模型实现网页数据的自动提取、去重和结构化处理。
- 从1.5万页网页中生成3531个独特习惯挑战,准确率高。
- 适合需要快速构建垂直领域数据集的研究者或开发者。
从海量非结构化网页中为特定应用构建专业内容语料库面临巨大数据整理挑战。本文提出一种高效流程,利用大语言模型(LLMs)实现网页数据的获取、筛选、结构化与清洗,显著提升数据整理规模与质量。我们设计了一套融合LLM增强技术的框架,用于结构化内容提取与语义去重。通过在行为教育领域集成到「30 Day Me」习惯养成应用中进行验证,所提出的30DayGen数据流水线从超过1.5万篇网页中提取并合成3531个独特的30天挑战。用户调查结果显示满意度达4.3/5,91%受访者表示愿意使用该内容实现习惯目标。
原文摘要 · Abstract (English)
Constructing specialized content corpora from vast, unstructured web sources for domain-specific applications poses substantial data curation challenges. In this paper, we introduce a streamlined approach for generating high-quality, domain-specific corpora by efficiently acquiring, filtering, structuring, and cleaning web-based data. We showcase how Large Language Models (LLMs) can be leveraged to address complex data curation at scale, and propose a strategical framework incorporating LLM-enhanced techniques for structured content extraction and semantic deduplication. We validate our approach in the behavior education domain through its integration into 30 Day Me, a habit formation application. Our data pipeline, named 30DayGen, enabled the extraction and synthesis of 3,531 unique 30-day challenges from over 15K webpages. A user survey reports a satisfaction score of 4.3 out of 5, with 91% of respondents indicating willingness to use the curated content for their habit-formation goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。