为大模型在企业真实场景中的评估建框架,解决学术测试与实际应用脱节问题。
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
- 构建十类企业关键任务评估体系,覆盖图像内容理解全链条。
- 提出BlockWeaver算法,高效比对无序文本输出,无需依赖嵌入或大模型。
- 基于7500张真实数据集,提供可落地的模型部署决策参考。
开源视觉语言模型在企业应用中前景广阔,但学术评估与企业部署需求存在显著差距。现有基准依赖多选题和合成数据,无法反映社交媒体内容分析等真实业务复杂性。本文提出VLM-in-the-Wild(ViLD)框架,从实际企业需求出发,定义十项关键任务:品牌标识检测、OCR、目标检测、人体存在与人口统计分析、人体行为与外观分析、场景识别、摄像头视角与媒体质量评估、主色调识别、综合描述生成及NSFW检测。为此,我们设计了创新的BlockWeaver算法,无需嵌入或大模型即可高效处理无序、分组不一的OCR输出,实现高可靠性和速度。为验证有效性,我们构建了包含7500个样本的新基准数据集,从百万级真实图像视频中严格分层采样。ViLD通过语义匹配(基于嵌入和大模型作为裁判)、传统指标及新方法,评估描述结果的完整性与忠实度。在主流开源模型(Qwen、MIMO、InternVL)与强大专有基线对比中,首次提供了以任务为导向、贴近工业实践的评估结果,为模型在企业环境中的部署提供切实可行的洞察。
原文摘要 · Abstract (English)
Open-source Vision-Language Models show immense promise for enterprise applications, yet a critical disconnect exists between academic evaluation and enterprise deployment requirements. Current benchmarks rely heavily on multiple-choice questions and synthetic data, failing to capture the complexity of real-world business applications like social media content analysis. This paper introduces VLM-in-the-Wild (ViLD), a comprehensive framework to bridge this gap by evaluating VLMs on operational enterprise requirements. We define ten business-critical tasks: logo detection, OCR, object detection, human presence and demographic analysis, human activity and appearance analysis, scene detection, camera perspective and media quality assessment, dominant colors, comprehensive description, and NSFW detection. To this framework, we bring an innovative BlockWeaver Algorithm that solves the challenging problem of comparing unordered, variably-grouped OCR outputs from VLMs without relying on embeddings or LLMs, achieving remarkable speed and reliability. To demonstrate efficacy of ViLD, we constructed a new benchmark dataset of 7,500 diverse samples, carefully stratified from a corpus of one million real-world images and videos. ViLD provides actionable insights by combining semantic matching (both embedding-based and LLM-as-a-judge approaches), traditional metrics, and novel methods to measure the completeness and faithfulness of descriptive outputs. By benchmarking leading open-source VLMs (Qwen, MIMO, and InternVL) against a powerful proprietary baseline as per ViLD framework, we provide one of the first industry-grounded, task-driven assessment of VLMs capabilities, offering actionable insights for their deployment in enterprise environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。