arXiv:2601.05609cs.CL2026-01中稿 · the Demonstration …被引 1

用大模型自动扩充法律文本数据,降低标注成本。

Data Augmented Pipeline for Legal Information Extraction and Reasoning

  • 用大模型生成法律信息抽取的增强数据
  • 减少人工标注量,提升系统鲁棒性
  • 方法通用,可拓展至其他NLP任务

本文提出一种基于大语言模型(LLMs)的法律领域信息抽取数据增强流水线。该方法简单有效,显著降低了人工标注所需的工作量,同时提升了信息抽取系统的鲁棒性。此外,该方法具备良好的泛化能力,可适用于法律领域之外的多种自然语言处理(NLP)任务。

原文摘要 · Abstract (English)

In this paper, we propose a pipeline leveraging Large Language Models (LLMs) for data augmentation in Information Extraction tasks within the legal domain. The proposed method is both simple and effective, significantly reducing the manual effort required for data annotation while enhancing the robustness of Information Extraction systems. Furthermore, the method is generalizable, making it applicable to various Natural Language Processing (NLP) tasks beyond the legal domain.

信息抽取法律AI数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。