arXiv:2607.20510cs.AIcs.CL2026-07

构建电信领域双语多模态评测基准,测试智能体跨源推理能力。

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

论文配图:Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
图 1 · 摘自论文原文
  • 设计多源异构数据融合的跨跳推理任务,支持中英文
  • 最强模型仅解决71%任务,视觉理解类低于30%
  • 适合评估企业级智能体,可复现且无需大模型裁判

我们提出Telco-GAIA,一个面向真实电信运营商数据的双语、多模态基准,用于评估工具使用型智能体。该基准包含100个经人工验证的问题-回答任务,涵盖英语和阿拉伯语,每项任务平均需要4.2步跨源推理,涉及三种异构数据源:静态网页快照(含HTML、图片和链接PDF)、合成关系型SQL数据库,以及外部网络存档,覆盖文本、图像和表格模态。基准以沙盒化Docker环境提供,通过归一化精确字符串匹配评分,确保评估客观、确定且可重复,无需依赖LLM作为评判者。在十二个商用和开源大模型上测试专用参考智能体,结果表明:即使最强模型也仅能完成71%的任务;在中等成本预算下,该比例降至约40%,其中视觉关联类任务表现最差,后端平均得分低于30%,文档与图像理解仍有巨大提升空间。Telco-GAIA为企事业单位智能体提供了严格可复现的测评平台,并为封闭领域基准建设提供了范例。

原文摘要 · Abstract (English)

We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human-verified question-answering tasks, in English and Arabic, that each demand multi-hop reasoning (4.2 hops on average) over three heterogeneous sources: a static website snapshot (HTML, images, and linked PDFs), a synthetic relational SQL database, and external web archives, spanning text, image, and tabular modalities. The benchmark is delivered as a sandboxed Docker environment and scored by normalized exact string matching, making evaluation objective, deterministic, and reproducible over time without any LLM-as-a-Judge. Evaluating a purpose-built reference agent across twelve commercial and open LLMs, we find Telco-GAIA challenging: even the strongest model solves only 71% of tasks; under a moderate cost budget, this falls to about 40%, and the visually grounded categories remain the weakest, where the average backend scores below 30%, leaving substantial headroom in document and image understanding. Telco-GAIA offers a rigorous, reproducible testbed for enterprise agents and a template for constructing closed-domain benchmarks.

智能体评测多模态电信领域双语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。