评测大模型跨语言网页搜索能力,聚焦英葡双语场景
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
- 设计双语检查清单,量化答案完整性和正确性
- 14个模型测试显示,编排机制提升信息覆盖度
- 英译葡存在显著性能波动,适合多语言研究者
大型语言模型日益作为信息源使用,但其可靠性取决于网络搜索、相关证据选取与答案整合能力。尽管已有基准评估网络浏览和代理工具使用,但多语言设置尤其是葡萄牙语仍研究不足。本文提出 extsc{MARCA},一个涵盖英语和葡萄牙语的双语基准,用于评估大模型在基于网页的信息检索能力。 extsc{MARCA} 包含52个由人工编写、涉及多个实体的问题,配有手动验证的检查清单式评分标准,明确衡量答案的完整性和正确性。我们在两种交互模式下评估14个模型:基础框架(直接搜索与抓取)与编排器框架(通过委托子代理实现任务分解)。为捕捉随机性,每个问题多次执行,并报告运行级别不确定性。结果显示,不同模型间表现差异显著,编排常提升覆盖度,且模型从英语向葡萄牙语迁移时存在明显性能波动。基准代码已公开于 https://github.com/maritaca-ai/MARCA。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesize complete answers. While recent benchmarks evaluate web-browsing and agentic tool use, multilingual settings, and Portuguese in particular, remain underexplored. We present \textsc{MARCA}, a bilingual (English and Portuguese) benchmark for evaluating LLMs on web-based information seeking. \textsc{MARCA} consists of 52 manually authored multi-entity questions, paired with manually validated checklist-style rubrics that explicitly measure answer completeness and correctness. We evaluate 14 models under two interaction settings: a Basic framework with direct web search and scraping, and an Orchestrator framework that enables task decomposition via delegated subagents. To capture stochasticity, each question is executed multiple times and performance is reported with run-level uncertainty. Across models, we observe large performance differences, find that orchestration often improves coverage, and identify substantial variability in how models transfer from English to Portuguese. The benchmark is available at https://github.com/maritaca-ai/MARCA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。