打造能实时对话的智能代理中介,让长任务更透明可控
Just A Rather Very Intelligent Spoken Agent

- 设计可实时交互的语音中介系统,持续连接用户与执行代理
- 在34个任务中验证,适时引入用户指导可提升任务完成率
- 中介的AI核心能力决定效果,适合研究人机协作与代理系统者
长时程AI代理日益强大,但与用户的互动仍显单薄。当前流程中,用户仅提供初始指令,后续仅接收片段文本更新,难以掌握代理进展或介入时机。为此,本文提出JarvisBench基准,用于评估中介在长任务流中的双重价值:一是协同任务完成度提升,二是用户交互的可理解性、响应性与可达性。我们构建了模块化参考原型,基于OpenClaw平台,在34个纯文本的WildClaw任务上评估了GPT-5.5、Claude Opus 4.7、Gemini及GPT系列代理。初步结果表明,具备追踪能力的语音中介可精准回答用户问题,并在适当时机注入稀疏引导时提升任务表现。效果高度依赖中介所用大模型,凸显该中间层的潜力与社区共建的必要性。
原文摘要 · Abstract (English)
Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin. In most workflows, users give an initial instruction, receive only selective textual updates, and lose a clear sense of what the agent is doing or when to step in. This leaves a missing part in the current agent ecosystem: an always-on Jarvis-style mediator that keeps the agent continuously reachable to the user. Such a mediator should support real-time spoken interaction with the user, answer questions without interrupting the worker, proactively report progress or confusion, and inject user guidance back into the agent's execution when useful. In this work, we introduce JarvisBench, a benchmark for measuring the dual value of mediation in long-horizon agent workflows. JarvisBench contains two complementary tracks: an agent-collaboration track that measures whether mediation improves downstream task completion, and a user-interaction track that measures whether mediation makes ongoing execution more understandable, responsive, and accessible to users. We instantiate the benchmark with a modular reference Jarvis prototype and evaluate it on 34 text-only WildClaw tasks executed in OpenClaw. Preliminary results with GPT-5.5, Claude Opus 4.7, Gemini-based, and GPT-based worker agents suggest that Jarvis-style mediation can provide trace-grounded responses to user questions and improve task performance when sparse user guidance is injected at appropriate moments. The results also show that effectiveness depends strongly on the mediator's LLM brain, highlighting both the promise of this missing middle layer and the need for broader community effort. Demo page https://cchen1436.github.io/jarvis
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。