arXiv:2606.07020cs.CL2026-06

MADE通过多语言智能体诊断引擎,实现细粒度模型评估洞察。

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

论文配图:MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights
图 1 · 摘自论文原文
  • 构建多语言诊断流程:规划、聚合分析、案例检查、文化反思与报告生成
  • 在34个文化、26种语言上,诊断报告质量比基线高47%
  • 适合多语言评估专家和模型迭代团队使用

多语言多文化评测已覆盖33个模型家族、11个基准、26种语言、34个文化,共866万条评估记录,但评分结果仍缺乏深入洞察。现有单一大模型或开放代理易被长而嘈杂的诊断输入淹没,且无通用诊断分类体系。为此,我们提出MADE——多语言智能体诊断引擎,将后评估分析分解为规划、聚合分析、实例级案例检查、多语言与文化反思、以及基于事实的报告合成五个步骤。MADE配套由专家设计的54个查询、15种语言的诊断集,在大规模多语言评测底座上验证。实验表明,MADE在诊断报告质量上比最强基线提升47%,人类多语言专家在87.9%的两两对比中更偏好它。结合多语言专家应用,MADE进一步揭示了部署、迭代与跨文化陷阱中的四项可行动发现,使评分表转化为模型选择与改进指南。

原文摘要 · Abstract (English)

Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and insight-poor, necessitating fine-grained multilingual post-evaluation diagnosis. However, single LLMs and open-ended agents are easily swamped by the long, noisy diagnostic input, and no reusable taxonomy exists for it. To address this, we propose MADE, a Multilingual Agentic Diagnosing Engine that decomposes post-evaluation analysis into planning, aggregate analysis, instance-level case inspection, multilingual and cultural reflection, and grounded report synthesis. MADE is paired with an expert-led 54-query and 15-language diagnostic set, evaluated on top of a large-scale multilingual evaluation substrate (33 model families, 11 benchmarks, 26 languages, 34 cultures, 8.66M evaluation records). Experiments show that MADE outperforms the strongest shared baseline by 47% in diagnosis report quality and is preferred by human multilingual experts in 87.9% of pairwise comparisons. Applied with multilingual experts, MADE further surfaces four actionable findings on deployment, iteration, and cross-cultural pitfalls, turning benchmark score tables into model-selection and remediation guidance.

多语言评估智能体诊断细粒度分析模型迭代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。