arXiv:2511.21750cs.CVcs.AI2025-11被引 5

首个评估多模态模型结构化输出能力的基准,覆盖四大视觉场景。

SO-Bench: A Structural Output Evaluation of Multimodal LLMs

  • 构建覆盖四类视觉场景的结构化输出评测基准
  • 发现主流模型在遵循数据模式方面仍有显著缺陷
  • 适合关注多模态推理与生成质量的研究者

多模态大语言模型(MLLMs)在实际代理场景中广泛应用,其输出不仅需准确,还需符合预定义的数据模式。尽管文本领域在结构化生成方面取得进展,但尚无系统性基准用于评估视觉输入下的模式约束信息抽取与推理。本文提出SO-Bench基准,涵盖用户界面、自然图像、文档和图表四个视觉领域,包含超过6500个多样化的JSON模式和1800对人工验证质量的图像-模式配对。对开源及前沿闭源模型的基准测试显示,模型在生成准确且符合模式的输出方面仍存在明显差距,凸显了改进多模态结构化推理的必要性。此外,我们通过训练实验显著提升了模型的结构化输出能力。基准与评估代码已公开于https://github.com/apple/ml-sobench。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are increasingly deployed in real-world, agentic settings where outputs must not only be correct, but also conform to predefined data schemas. Despite recent progress in structured generation in textual domain, there is still no benchmark that systematically evaluates schema-grounded information extraction and reasoning over visual inputs. In this work, we conduct a comprehensive study of visual structural output capabilities for MLLMs with our carefully designed SO-Bench benchmark. Covering four visual domains, including UI screens, natural images, documents, and charts, SO-Bench is built from over 6.5K diverse JSON schemas and 1.8K curated image-schema pairs with human-verified quality. Benchmarking experiments on open-sourced and frontier proprietary models reveal persistent gaps in predicting accurate, schema compliant outputs, highlighting the need for better multimodal structured reasoning. Beyond benchmarking, we further conduct training experiments to largely improve the model's structured output capability. We make the benchmark and evaluation publicly available at https://github.com/apple/ml-sobench

多模态结构化输出评测基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。