arXiv:2604.27137cs.CL2026-04

用语言能力标准评估Claude多语言回应一致性,发现差异可解释且影响公平部署。

Cross-Lingual Response Consistency in Large Language Models: An ILR-Informed Evaluation of Claude Across Six Languages

  • 基于ILR语言能力标准设计跨语言评估框架,覆盖六种语言。
  • 法语回复比德语长约30%,创意类内容跨语言差异最大。
  • 专家分析揭示五类语言特异性输出模式,适合多语言AI开发者参考。

本文提出一种基于国际语言培训委员会(ILR)技能等级描述的系统性评估框架,应用于Claude(Sonnet 4.6)在英语、法语、罗马尼亚语、西班牙语、意大利语和德语六种语言上的表现。通过12组语义等价提示,覆盖ILR 1至3+级复杂度,收集216条响应(12提示×6语言×3次运行),采用量化指标与专家质性评估双层方法分析。量化结果显示,相同提示下法语回复长度约为德语的1.3倍,创意与情感类提示组的跨语言表面差异最高。由具备12年ILR/OPI评估经验的六语种专业人员进行的质性分析识别出五类跨语言变异模式:语用消歧策略的系统性差异、创意输出中审美与文学传统的分歧、语言内技术术语规范差异、文化校准缺失导致缺乏特定文化内容而倾向中性模板、以及情感支持响应中的机构指派行为语言特异性。研究认为,基于ILR的专家判断为大模型输出提供了新颖且未被充分重视的评估路径,能补充纯计算基准,并表明Claude的跨语言输出差异具有可解释性、领域依赖性和对公平多语言部署的实质性影响。

原文摘要 · Abstract (English)

This paper introduces a systematic evaluation framework grounded in the Interagency Language Roundtable (ILR) Skill Level Descriptions and applies it to Claude (Sonnet 4.6) across six languages: English, French, Romanian, Spanish, Italian, and German. We administer a battery of 12 semantically equivalent prompt clusters spanning ILR complexity levels 1 through 3+, collect 216 responses (12 prompts, 6 languages, 3 runs), and analyze outputs through a two-layer methodology combining automated quantitative metrics with expert ILR qualitative assessment. Quantitative analysis reveals that French responses are approximately 30% longer than German responses on identical prompts, and that creative and affective clusters show the highest cross-lingual surface divergence. Qualitative analysis, conducted by a six-language professional with 12 years of ILR/OPI assessment experience, identifies five cross-lingual variation patterns: systematic differences in pragmatic disambiguation strategies, aesthetic and literary tradition divergence in creative output, language-internal technical terminology norms, cultural calibration gaps evidenced by the absence of culture-specific content in favor of culturally neutralized templates, and language-specific institutional referral behavior in emotional support responses. We argue that ILR-informed expert judgment applied to LLM outputs constitutes a novel and underreported evaluation methodology that complements purely computational benchmarks, and that cross-lingual output variation in Claude is interpretable, domain-dependent, and consequential for equitable multilingual AI deployment.

多语言大模型评估语言能力标准Claude

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。