通过可视化嵌入空间,定位文档理解模型的错误高发区并生成增强数据。
VERSE: Visual Embedding Reduction and Space Exploration. Clustering-Guided Insights for Training Data Enhancement in Visually-Rich Document Understanding
- 用聚类分析视觉嵌入空间,定位模型薄弱区域。
- 在合成数据上训练后,F1指标显著提升且泛化能力不降。
- 适配本地模型可媲美甚至超越云端大模型,适合资源受限场景。
本文提出VERSE方法,用于分析和改进视觉-语言模型在视觉丰富文档理解中的表现,通过探索其视觉嵌入空间实现。该方法支持可视化潜在表示,辅助评估模型可行性,并识别问题区域,指导生成合成数据以提升性能。我们在合成的MERIT数据集上训练,于真实世界数据集MERIT Secret上评估。结果表明,VERSE能揭示易错聚类对应的视觉特征,使用含这些特征的样本重新训练后,F1性能显著提升而泛化能力不受影响。此外,经VERSE优化的本地模型(如Donut、Idefics2)性能可达到甚至超过GPT-4和Pixtral等SaaS解决方案。
原文摘要 · Abstract (English)
This work introduces VERSE, a methodology for analyzing and improving Vision-Language Models applied to Visually-rich Document Understanding by exploring their visual embedding space. VERSE enables the visualization of latent representations, supporting the assessment of model feasibility. It also facilitates the identification of problematic regions and guides the generation of synthetic data to enhance performance in those clusters. We validate the methodology by training on the synthetic MERIT Dataset and evaluating on its real-world counterpart, MERIT Secret. Results show that VERSE helps uncover the visual features associated with error-prone clusters, and that retraining with samples containing these features substantially boosts F1 performance without degrading generalization. Furthermore, we demonstrate that on-premise models such as Donut and Idefics2, when optimized with VERSE, match or even surpass the performance of SaaS solutions like GPT-4 and Pixtral.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。