提出一套让视觉智能体又快又稳的架构设计方法。
A Pattern Language for Resilient Visual Agents

- 用快慢分离策略,区分实时反应与缓慢判断
- 四种模式协同,提升视觉代理在企业场景的可靠性
- 适合需要高实时性与稳定性的工业视觉系统开发者
将多模态基础模型集成到企业生态中面临根本性的软件架构挑战。架构师需权衡相互冲突的质量属性:视觉语言动作(VLA)模型的高延迟和非确定性,与企业控制回路对严格确定性和实时性能的要求。本研究提出一种视觉代理的架构模式语言,将快速确定性反射与慢速概率性监督相分离。该语言包含四个架构设计模式:(1) 混合可操作性融合,(2) 自适应视觉锚定,(3) 视觉层次合成,(4) 语义场景图。
原文摘要 · Abstract (English)
Integrating multimodal foundation models into enterprise ecosystems presents a fundamental software architecture challenge. Architects must balance competing quality attributes: the high latency and non-determinism of vision language action (VLA) models versus the strict determinism and real-time performance required by enterprise control loops. In this study, we propose an architectural pattern language for visual agents that separates fast, deterministic reflexes from slow, probabilistic supervision. It consists of four architectural design patterns: (1) Hybrid Affordance Integration, (2) Adaptive Visual Anchoring, (3) Visual Hierarchy Synthesis, and (4) Semantic Scene Graph.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。