arXiv:2412.13859cs.CV2024-12被引 10

用大模型零样本提示与少量微调,大幅减少文档分类标注数据需求。

Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models

  • 基于大模型的零样本提示和少样本微调策略
  • 在少量甚至无标注数据下实现接近顶尖性能
  • 适合低资源场景下的文档理解应用

对扫描文档进行分类是一项挑战性任务,涉及图像、版式和文本分析以实现文档理解。然而,在某些基准数据集(如RVL-CDIP)上,当使用数十万训练样本时,现有方法已接近完美性能。随着大型语言模型(LLMs)作为出色的少样本学习者出现,问题在于:仅用少量甚至零标注样本,能否解决文档分类问题?本文在零样本提示与少样本微调背景下探讨此问题,旨在尽可能减少对人工标注训练样本的依赖。

原文摘要 · Abstract (English)

Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is closing in to near-perfect performance when considering hundreds of thousands of training samples. With the advent of large language models (LLMs), which are excellent few-shot learners, the question arises to what extent the document classification problem can be addressed with only a few training samples, or even none at all. In this paper, we investigate this question in the context of zero-shot prompting and few-shot model fine-tuning, with the aim of reducing the need for human-annotated training samples as much as possible.

文档分类大模型少样本学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。