arXiv:2607.06008cs.AIcs.CL2026-07

评测大模型在跨语言长流程任务中的表现,发现多语言让执行失败率飙升。

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows

论文配图:PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
图 1 · 摘自论文原文
  • 构建跨语言长周期工作流基准,含67个真实场景任务
  • 多语言任务下模型性能显著下降,尤其在复杂流程中
  • 适合关注多语言AI代理、企业自动化落地的研究者

尽管大语言模型代理在单语种长周期规划与工具使用方面表现优异,但企业工作流天然需要处理跨语言资源并完成长期任务。然而,多语言与长周期执行之间的交互仍缺乏研究。本文提出PolyWorkBench,一个用于评估大模型代理在跨语言长周期工作任务中表现的基准。该基准涵盖商业、知识工作、法律分析、本地化和制造五个核心领域,共67项任务,均由作者基于真实数据生成,并经第二作者独立审核验证。代理需整合异构多语言输入,执行迭代式工具使用路径,并生成结构化领域输出。为严谨评估,采用任务特定的结构评分标准Grade作为主指标,辅以Pytest进行可执行状态验证,及LLM-as-Judge进行语义质量诊断。评估结果表明,代理在不同语言间表现差异显著,且在更复杂的跨语言任务中性能急剧下降;分析显示,多语言执行暴露了规划、工具交互与决策环节中的系统性失败模式。

原文摘要 · Abstract (English)

While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored. We introduce PolyWorkBench, a benchmark designed to evaluate LLM agents on multilingual, long-horizon workplace workflows. PolyWorkBench features 67 tasks across five core domains: commerce, knowledge work, legal analysis, localization, and manufacturing. Tasks are authored by the paper's authors from real-world data seeds and independently verified through a second-author audit. Agents must integrate heterogeneous multilingual inputs, execute iterative tool-use trajectories, and produce structured domain artifacts. To rigorously assess performance, we adopt Grade, a task-specific structural scoring rubric, as our primary ranking metric, and complement it with Pytest for executable state verification and LLM-as-Judge for semantic quality diagnostics. Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and our analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.

大模型代理跨语言长周期任务基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。