arXiv:2608.03735cs.MAcs.CL2026-08

针对多语言智能体规划失败,提出可操作诊断框架并提升跨语言任务准确率。

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

论文配图:An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures
图 1 · 摘自论文原文
  • 基于真实任务失败案例,构建多语言规划失误的可操作分类体系。
  • 在低资源语言中,规划错误占比显著上升,最高达70%以上。
  • 引入TART框架后,多语言任务平均准确率提升5.6个百分点。

多语言多智能体系统在非英语语境下性能显著下降,但以往研究极少揭示用户请求转化为可执行计划过程中关键信息的丢失机制。本文将规划器视为请求到动作的接口,通过分析真实任务失败案例,构建了可操作的规划对齐失败分类体系。基于大模型的分析表明,随着语言资源减少,规划失败在未成功执行中占比持续上升,低资源语言中影响最显著。为验证该分类体系的可缓解性,本文提出TART(Taxonomy-Guided Actionable Representation)框架,使分类关键要素显式暴露于规划器及下游子智能体。在多种语言、三种LLM骨干网络、两个数据集和两种代理配置下,TART均稳定提升性能。在涵盖11种语言的多语言GAIA基准上,其将先进系统平均准确率提升5.6个百分点。

原文摘要 · Abstract (English)

Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is lost when user requests are converted into executable plans. We study the planner in a multi-agent system as the request-to-action interface and derive an actionable taxonomy of planning-grounding failures from failed real-world task executions. LLM-based analysis shows that these failures constitute an increasing share of unsuccessful executions as language-resource availability declines, with the strongest effects in low-resource languages. To test whether the taxonomy supports mitigation, we introduce TART, Taxonomy-Guided Actionable Representation, that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents. Across multiple languages, three LLM backbones, two datasets, and two agentic configurations, TART consistently improves performance. On multilingual GAIA, it raises a state-of-the-art system's accuracy by 5.6 percentage points averaged across eleven languages spanning low- to high-resource settings.

多语言智能体规划诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。