arXiv:2511.18335cs.CLcs.AI2025-11

让小模型通过合成数据训练,实现跨多种结构的文本转结构生成。

OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas

  • 用统一任务框架整合多类结构生成任务
  • 仅用合成数据微调小模型,性能接近GPT-4o
  • 为文本转结构提供可复现的评估基准

大型语言模型(LLMs)生成符合任意模式的结构化输出的能力,对信息抽取、表格生成和函数调用等下游任务至关重要。尽管现代LLMs在自然语言生成方面表现优异,但其在文本转结构任务上的能力仍不明确。为此,我们提出OmniStruct,一个涵盖信息抽取、表格生成和函数调用等多种文本转结构任务的综合性基准。该基准通过整合现有适配结构化答案格式的数据集,并将其统一到一致的任务设置下构建而成。为促进高效文本转结构模型的发展,我们通过合成任务生成高质量训练数据。实验表明,在未使用任何监督数据的情况下,仅用合成数据微调的小模型即可成为通用结构生成模型,性能媲美GPT-4o。

原文摘要 · Abstract (English)

The ability of Large Language Models (LLMs) to generate structured outputs that follow arbitrary schemas is crucial to a wide range of downstream tasks that require diverse structured representations of results such as information extraction, table generation, and function calling. While modern LLMs excel in generating unstructured responses in natural language, whether this advancement translates to a strong performance on text-to-structure tasks remains unclear. To bridge this gap, we first introduce OmniStruct, a comprehensive benchmark for assessing LLMs' capabilities on diverse text-to-structure tasks such as information extraction, table generation, and function calling. We build OmniStruct by identifying existing datasets across a wide range of tasks that are suitable for a structured answer format, and adapting them under a unified text-to-structure problem setting. To facilitate the development of efficient text-to-structure models, we collect high-quality training data via synthetic task generation. Without using any supervised data for OmniStruct tasks, our experiments demonstrate the possibility of fine-tuning much smaller models on synthetic data into universal structured generation models that can rival the performance of GPT-4o.

文本生成结构化输出小模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。