大模型可精准将菜谱文本转为标准格式,无需大量训练
The Effectiveness of Large Language Models in Transforming Unstructured Text to Standardized Formats
- 用少样本提示让大模型直接转换非结构化菜谱为标准化格式
- GPT-4o在转换中达到ROUGE-L 0.9722、WER 0.0730的高精度
- 小模型经微调也能高效完成任务,适合资源有限场景
非结构化文本数据的爆炸式增长给现代数据管理与信息检索带来根本挑战。尽管大语言模型(LLMs)在自然语言处理中表现出色,其将非结构化文本转化为标准化结构化格式的潜力仍基本未被探索——这一能力有望彻底改变各行业的数据处理流程。本研究首次系统评估了LLMs将非结构化菜谱文本转化为结构化Cooklang格式的能力。通过测试四种模型(GPT-4o、GPT-4o-mini、Llama3.1:70b、Llama3.1:8b),引入结合传统指标(WER、ROUGE-L、TER)与语义元素识别专用指标的创新评估方法。实验表明,采用少样本提示的GPT-4o实现突破性表现(ROUGE-L: 0.9722,WER: 0.0730),首次证明大模型可在无需大量训练的情况下可靠完成特定领域文本的结构化转换。尽管模型性能总体随规模提升,但发现如Llama3.1:8b等小型模型通过针对性微调亦具显著潜力。这些发现为医疗记录、技术文档等多领域自动化结构化数据生成开辟新路径,可能重塑组织处理与利用非结构化信息的方式。
原文摘要 · Abstract (English)
The exponential growth of unstructured text data presents a fundamental challenge in modern data management and information retrieval. While Large Language Models (LLMs) have shown remarkable capabilities in natural language processing, their potential to transform unstructured text into standardized, structured formats remains largely unexplored - a capability that could revolutionize data processing workflows across industries. This study breaks new ground by systematically evaluating LLMs' ability to convert unstructured recipe text into the structured Cooklang format. Through comprehensive testing of four models (GPT-4o, GPT-4o-mini, Llama3.1:70b, and Llama3.1:8b), an innovative evaluation approach is introduced that combines traditional metrics (WER, ROUGE-L, TER) with specialized metrics for semantic element identification. Our experiments reveal that GPT-4o with few-shot prompting achieves breakthrough performance (ROUGE-L: 0.9722, WER: 0.0730), demonstrating for the first time that LLMs can reliably transform domain-specific unstructured text into structured formats without extensive training. Although model performance generally scales with size, we uncover surprising potential in smaller models like Llama3.1:8b for optimization through targeted fine-tuning. These findings open new possibilities for automated structured data generation across various domains, from medical records to technical documentation, potentially transforming the way organizations process and utilize unstructured information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。