小样本下高效适配文档理解的图网络模型
FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
- 模块化设计融合语言与视觉骨干,支持少样本快速适应
- 参数少于90万,收敛快且在真实场景中抗错能力强
- 适合资源受限环境下的信息抽取应用
本文提出少样本域自适应图网络(FS-DAG),一种适用于视觉丰富文档理解(VRDU)的可扩展、高效模型架构。该模型在模块化框架内结合领域特定与语言/视觉专用主干网络,仅用少量数据即可适应多种文档类型。其对实际部署中的光学字符识别错误、拼写错误和域偏移具有鲁棒性。FS-DAG 参数量低于90M,性能优异,特别适合计算资源有限的信息抽取任务。通过大量实验验证,相比现有先进方法,其在收敛速度与准确率上均有显著提升。本工作也推动了小型化、高效化模型的发展,兼顾性能与实用性。
原文摘要 · Abstract (English)
In this work, we propose Few Shot Domain Adapting Graph (FS-DAG), a scalable and efficient model architecture for visually rich document understanding (VRDU) in few-shot settings. FS-DAG leverages domain-specific and language/vision specific backbones within a modular framework to adapt to diverse document types with minimal data. The model is robust to practical challenges such as handling OCR errors, misspellings, and domain shifts, which are critical in real-world deployments. FS-DAG is highly performant with less than 90M parameters, making it well-suited for complex real-world applications for Information Extraction (IE) tasks where computational resources are limited. We demonstrate FS-DAG's capability through extensive experiments for information extraction task, showing significant improvements in convergence speed and performance compared to state-of-the-art methods. Additionally, this work highlights the ongoing progress in developing smaller, more efficient models that do not compromise on performance. Code : https://github.com/oracle-samples/fs-dag
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。