arXiv:2608.05864cs.AI2026-08

多模态大模型在决策中感知视觉信息,但可能因信息过载导致资源分配失误。

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

论文配图:Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
图 1 · 摘自论文原文
  • 构建多模态企业决策评测基准,对比文本与图文结合的决策表现。
  • 视觉信息提升风险预测和解释能力,但削弱资源分配准确性。
  • 视觉信号过多引发干扰,建议有选择地引入视觉输入。

大型语言模型正被用于自主决策代理,但现有企业决策评估多局限于纯文本场景,无法判断模型是否能有效利用视觉商业证据。为此,我们提出 C-SUITEBENCH,一个受控的多模态基准,包含50个场景下的5类决策任务,在文本与图文双条件下进行对比。将9个前沿模型置于首席执行官角色进行评估。结果显示,多模态输入显著提升以证据为中心的推理能力,尤其在风险预测与董事会陈述方面收益最大。然而,我们发现“多模态整合悖论”:尽管视觉信息增强了对事实的把握,但所有模型在受限资源分配上表现反而下降。消融实验表明,这是由于信号拥挤所致——各视觉通道单独有效,但组合使用会干扰解码过程中的约束满足。研究揭示视觉感知与受约束行动是独立瓶颈,盲目增加视觉输入可能损害高风险决策,提示未来企业级AI应采用选择性视觉融合策略。

原文摘要 · Abstract (English)

Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.

多模态决策大模型应用企业AI视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。