Orion用工具链实现多步视觉智能,让机器能像人一样看图做事。
Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- 用物体检测、OCR等工具链执行复杂视觉任务
- 在多个评测中表现优异,超越传统模型
- 适合需要自主决策的视觉应用开发
我们提出Orion,一个将视觉推理与工具增强执行结合的视觉智能体,可处理图像、视频和文档中的多步视觉任务。不同于生成描述的传统视觉语言模型,Orion协同使用目标检测、关键点定位、全景分割、光学字符识别(OCR)和几何分析等专用计算机视觉工具,完成复杂的多步视觉工作流。系统在MMMU、MMBench、DocVQA和MMLongBench等多个基准上达到竞争力表现,将单体视觉语言模型能力拓展至生产级视觉智能水平。通过代理式、工具增强的方法,Orion实现了从被动感知到主动工具驱动的视觉智能跃迁,标志着视觉理解向可执行智能的演进。免费体验:https://chat.vlm.run,更多信息:https://www.vlm.run/orion
原文摘要 · Abstract (English)
We introduce Orion, a visual agent that integrates vision-based reasoning with tool-augmented execution to achieve powerful, precise, multi-step visual intelligence across images, video, and documents. Unlike traditional vision-language models that generate descriptive outputs, Orion orchestrates a suite of specialized computer vision tools, including object detection, keypoint localization, panoptic segmentation, Optical Character Recognition (OCR), and geometric analysis, to execute complex multi-step visual workflows. The system achieves competitive performance across MMMU, MMBench, DocVQA, and MMLongBench while extending monolithic VLM capabilities to production-grade visual intelligence. Through its agentic, tool-augmented approach, Orion enables autonomous visual reasoning that bridges neural perception with symbolic execution, marking the transition from passive visual understanding to active, tool-driven visual intelligence. Try Orion for free at: https://chat.vlm.run Learn more at: https://www.vlm.run/orion
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。