量化多模态输入对企业文档理解的信任度影响,提升系统可靠性。
Evaluating VisualRAG: Quantifying Cross-Modal Performance in Enterprise Document Understanding
- 构建跨模态输入的量化评估框架,分析文本、图像、字幕、OCR的权重影响。
- 最优权重组合使性能比纯文本提升57.3%,兼顾效率与准确率。
- 揭示大模型在生成字幕和提取文字中的差异,适合关注可信AI的企业用户。
当前多模态生成式AI的评估框架难以建立可信度,阻碍了企业应用中对可靠性的需求。本文提出一种系统性、量化的基准测试框架,用于衡量VisualRAG系统在企业文档智能中逐步融合文本、图像、字幕和OCR等跨模态输入时的可信度。该方法建立了技术指标与用户信任度之间的定量关系。实验表明,采用30%文本、15%图像、25%字幕、30%OCR的最优权重组合,可使性能相比纯文本基线提升57.3%,同时保持计算高效。我们还对比评估了基础模型在字幕生成与OCR提取中的表现差异,这对企业级可信AI部署至关重要。本工作通过严谨框架推动多模态RAG在关键企业场景中的负责任落地。
原文摘要 · Abstract (English)
Current evaluation frameworks for multimodal generative AI struggle to establish trustworthiness, hindering enterprise adoption where reliability is paramount. We introduce a systematic, quantitative benchmarking framework to measure the trustworthiness of progressively integrating cross-modal inputs such as text, images, captions, and OCR within VisualRAG systems for enterprise document intelligence. Our approach establishes quantitative relationships between technical metrics and user-centric trust measures. Evaluation reveals that optimal modality weighting with weights of 30% text, 15% image, 25% caption, and 30% OCR improves performance by 57.3% over text-only baselines while maintaining computational efficiency. We provide comparative assessments of foundation models, demonstrating their differential impact on trustworthiness in caption generation and OCR extraction-a vital consideration for reliable enterprise AI. This work advances responsible AI deployment by providing a rigorous framework for quantifying and enhancing trustworthiness in multimodal RAG for critical enterprise applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。