arXiv:2509.08381cs.CLcs.AI2025-09

小样本下用百亿参数模型实现多任务结构化信息抽取

Low-Resource Fine-Tuning for Multi-Task Structured Information Extraction with a Billion-Parameter Instruction-Tuned Model

  • 用低秩适配微调百亿参数模型,每任务仅需数百到千个样本
  • 在最低数据量下仍显著优于基线,多数指标表现更优
  • 适合计算资源有限但需高可靠性的企业级信息抽取场景

将大语言模型应用于金融合规报告、法律文档分析及多语言知识库构建等结构化数据提取任务时,小团队常因运行大型模型成本过高且难以获取高质量数据集而受限。现有指令微调研究多聚焦于七亿参数及以上模型,缺乏对小型模型在低资源、多任务条件下有效性的验证。本文提出ETLCH,基于LLaMA的百亿参数模型,采用低秩适配技术,在每项任务仅使用数百至千个样本的情况下,完成JSON提取、知识图谱抽取和命名实体识别。尽管规模较小,ETLCH在多数评估指标上超越强基线,尤其在最低数据量下仍展现显著优势。结果表明,经过良好调优的小型模型可在极低计算成本下提供稳定准确的结构化输出,为资源受限环境下的高效可靠信息抽取提供了可行方案。

原文摘要 · Abstract (English)

Deploying large language models (LLMs) for structured data extraction in domains such as financial compliance reporting, legal document analytics, and multilingual knowledge base construction is often impractical for smaller teams due to the high cost of running large architectures and the difficulty of preparing large, high-quality datasets. Most recent instruction-tuning studies focus on seven-billion-parameter or larger models, leaving limited evidence on whether much smaller models can work reliably under low-resource, multi-task conditions. This work presents ETLCH, a billion-parameter LLaMA-based model fine-tuned with low-rank adaptation on only a few hundred to one thousand samples per task for JSON extraction, knowledge graph extraction, and named entity recognition. Despite its small scale, ETLCH outperforms strong baselines across most evaluation metrics, with substantial gains observed even at the lowest data scale. These findings demonstrate that well-tuned small models can deliver stable and accurate structured outputs at a fraction of the computational cost, enabling cost-effective and reliable information extraction pipelines in resource-constrained environments.

信息抽取小样本学习低资源大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。