arXiv:2603.22576cs.CL2026-03

用巴西文学作品测试大模型指令遵循能力,更真实可靠。

CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context

  • 以八部巴西文学经典为背景设计可自动验证的指令任务
  • 顶级模型严格准确率达98.5%,专用模型成本低至0.13美元
  • 发现形态约束和多轮对话中指令保持是主要挑战

我们提出CAPITU,一个评估大语言模型在巴西葡萄牙语中指令遵循能力的基准。不同于聚焦英语或使用通用提示的现有基准,CAPITU将所有任务置于八部巴西文学经典背景下,结合可验证的指令约束与文化相关的内容。该基准包含59种指令类型,分为七个类别,均支持自动验证,无需依赖LLM评判或人工评估。指令类型涵盖葡萄牙语特有的语言约束(如-ando/-endo/-indo、-inho/-inha、-mente词尾)和结构要求。我们在单轮和多轮设置下评估了18个前沿模型。结果表明,顶尖推理模型表现优异(GPT-5.2带推理:98.5%严格准确率),而专用于葡萄牙语的模型则具备更高性价比(Sabiazinho-4:87.0%准确率,耗资0.13美元;Claude-Haiku-4.5:73.5%准确率,耗资1.12美元)。多轮评估显示指令持续性差异显著,对话级准确率在60%至96%之间波动。研究识别出形态约束、精确计数及多轮中约束退化等具体挑战。我们已公开完整基准、评估代码和基线结果,以推动葡萄牙语指令遵循研究。

原文摘要 · Abstract (English)

We introduce CAPITU, a benchmark for evaluating instruction-following capabilities of Large Language Models (LLMs) in Brazilian Portuguese. Unlike existing benchmarks that focus on English or use generic prompts, CAPITU contextualizes all tasks within eight canonical works of Brazilian literature, combining verifiable instruction constraints with culturally-grounded content. The benchmark comprises 59 instruction types organized into seven categories, all designed to be automatically verifiable without requiring LLM judges or human evaluation. Instruction types include Portuguese-specific linguistic constraints (word termination patterns like -ando/-endo/-indo, -inho/-inha, -mente) and structural requirements. We evaluate 18 state-of-the-art models across single-turn and multi-turn settings. Our results show that frontier reasoning models achieve strong performance (GPT-5.2 with reasoning: 98.5% strict accuracy), while Portuguese-specialized models offer competitive cost-efficiency (Sabiazinho-4: 87.0% at \$0.13 vs Claude-Haiku-4.5: 73.5% at \$1.12). Multi-turn evaluation reveals significant variation in constraint persistence, with conversation-level accuracy ranging from 60% to 96% across models. We identify specific challenges in morphological constraints, exact counting, and constraint persistence degradation across turns. We release the complete benchmark, evaluation code, and baseline results to facilitate research on instruction-following in Portuguese.

指令遵循巴西葡语多轮对话评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。