用智能代理预测多语言模型在缺数据时的表现,提升跨语言部署可靠性。
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
- 构建1500题基准测试,分离可得证据与真实答案,模拟真实评估场景。
- 提出基于有向无环图的智能体系统,在无直接数据时仍能准确预测模型表现。
- 适合需要跨语言部署、缺乏评测数据的研究者和工程团队使用。
我们研究预测性多语言评估:当目标语言的任务评测结果缺失时,如何估计模型的表现。该问题在多语言部署中普遍存在,因评测覆盖不均、文献分布不均导致。我们设计了一个包含1,500个问题的可控基准,涵盖六项任务与五种证据场景,将可获取证据与真实答案分离,以评估需从不完整文献中推断结果的系统。同时提出Litmus (Re)Agent,一种基于有向无环图(DAG)调度的智能体系统,将查询分解为假设,检索证据,并通过特征感知聚合生成预测。在六种系统中,Litmus (Re)Agent整体表现最佳,尤其在直接证据薄弱或缺失的迁移密集场景中提升显著。结果表明,结构化智能体推理是应对不完整证据下多语言性能估计的有效方法。
原文摘要 · Abstract (English)
We study predictive multilingual evaluation: estimating how well a model will perform on a task in a target language when direct benchmark results are missing. This problem is common in multilingual deployment, where evaluation coverage is sparse and published evidence is uneven across languages, tasks, and model families. We introduce a controlled benchmark of 1,500 questions spanning six tasks and five evidence scenarios. The benchmark separates accessible evidence from ground truth, enabling evaluation of systems that must infer missing results from incomplete literature evidence. We also present Litmus (Re)Agent, a DAG-orchestrated agentic system that decomposes queries into hypotheses, retrieves evidence, and synthesises predictions through feature-aware aggregation. Across six systems, Litmus (Re)Agent achieves the best overall performance, with the largest gains in transfer-heavy scenarios where direct evidence is weak or absent. These results show that structured agentic reasoning is a promising approach to multilingual performance estimation under incomplete evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。