arXiv:2601.16217cs.CLcs.AI2026-01

评测大模型在中英术语混用场景下的表现,发现模型偏好与专业习惯有差距。

ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing

  • 构建可控基准测试集,聚焦中文技术语境中的中英术语混合现象
  • 9个模型在1167对样本中更倾向使用中文术语,但专业术语偏好不稳健
  • 提供可复用的诊断工具,适合研究多语言社区惯例的学者使用

大型语言模型日益成为多语言专业交流的中介,其有效生成需适应社区对术语保留、翻译或混用的习惯。现有基准很少能独立考察此类受社区约束的选择。本文提出ChiEngMixBench,一个针对中文人工智能/计算机科学语境的控制型基准,其中中文句式常嵌入已确立的英文技术术语。该数据集源自公开技术讨论,包含1,706个源文本衍生的候选对,涵盖1,344个非空标准化术语,其中1,167对为严格子集:固定中文前缀和句法位置,仅改变术语形式。基准结合成对似然比较与透明参考轮廓诊断,用于开放生成响应。在九个开源模型中,中文等价词在多数配对中获得更高似然,揭示源文本验证用法与模型偏好之间的差距。专业术语虽有微弱方向性提升,但在控制频率、长度及多重比较校正后并不显著。人工评估与基线分析表明,参考轮廓一致性在预设混用风格下具信息量,但无法可靠预测整体偏好。ChiEngMixBench为社区特异性多语言惯例研究提供了可复用的测试平台,并设定了明确诊断边界。

原文摘要 · Abstract (English)

Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community conventions about which expressions are retained, translated, or mixed. Existing benchmarks rarely isolate such community-conditioned choices. We introduce ChiEngMixBench, a controlled benchmark for Chinese AI/CS discourse, where Chinese frames routinely incorporate established English technical terms. Built from public technical discussions, it contains 1,706 source-derived candidate pairs covering 1,344 non-empty normalized terms, including a 1,167-pair strict subset that fixes the Chinese prefix and syntactic position while varying only the terminology form. The benchmark combines paired likelihood comparisons with a transparent reference-profile diagnostic for open-ended responses. Across nine open-weight models, Chinese equivalents receive higher likelihood on most pairs, revealing a gap between source-attested usage and model preference. Specialized terms show a small directional lift that is not robust after frequency and length controls and multiple-comparison correction. Human evaluation and baseline analyses show that reference-profile conformity is informative under the intended mixed-style rubric but does not reliably predict holistic response preference. ChiEngMixBench provides a reusable testbed for community-specific multilingual conventions with explicit diagnostic boundaries.

多语言术语混用评测基准AI写作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。