arXiv:2604.27419cs.AIcs.CL2026-04被引 1

首个模拟非专家用户交互的网页生成评测,揭示大模型在模糊指令下的执行困境。

InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?

论文配图:InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
图 1 · 摘自论文原文
  • 设计四类用户代理和扰动策略,模拟真实低代码场景中的指令混乱
  • 构建包含澄清、实现、验证等动作的交互环境,支持迭代优化
  • 发现前沿模型仍陷于盲执行,暴露其意图理解与自适应能力短板

随着多模态大语言模型(MLLM)和编程智能体的发展,网页开发正从手动编码转向基于智能体的项目级代码合成。现有评测依赖理想化假设,如结构良好、信息丰富的输入和静态执行环境。而现实开发面临关键瓶颈:非专家用户提供的模糊、低质量指令与模型理解间存在语义错位,导致我们称之为‘盲执行’的失败模式。为填补这一空白,我们提出 InteractWeb-Bench,首个在非专家低代码用户条件下评估网页生成的多模态交互评测基准。该基准引入四类用户代理及基于需求工程缺陷分类体系的情境扰动,系统模拟指令模糊、冗余、矛盾等多样化用户行为。我们构建了支持澄清、实现、验证、提交统一动作空间的交互执行环境,实现意图迭代修正、代码生成与视觉反馈验证。大量实验与分析表明,前沿的 MLLM 智能体仍困于盲执行,暴露出其在意图识别与自适应交互方面的局限性。

原文摘要 · Abstract (English)

With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code synthesis. Existing benchmarks rely on idealized assumptions, especially for well-structured, information-rich inputs and static execution settings. In contrast, real-world development is constrained by a critical bottleneck: the semantic misalignment between ambiguous, low-quality instructions from non-expert users and model understanding, which results in a failure mode that we term blind execution. To address this gap, we introduce InteractWeb-Bench, the first multimodal interactive benchmark for website generation under non-expert low-code user conditions. InteractWeb-Bench introduces four types of user agents and persona-driven instruction perturbations to systematically simulate diverse user behaviors, including ambiguity, redundancy, and contradiction, grounded in requirement engineering defect taxonomies. We develop an interactive execution environment for agents, featuring a unified action space comprising Clarify, Implement, Verify, and Submit, enabling iterative intent refinement, code synthesis, and visual feedback-based validation. Extensive experiments and analysis reveal that frontier MLLM-based agents remain trapped in blind execution, exposing limitations in intent recognition and adaptive interaction.

多模态智能体网页生成评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。