构建首个印地语罗马化混合指令评测基准,评估大模型在多语言混用场景下的表现。
Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions

- 设计七类任务、四种语言、三档混写强度的罗马化混合指令数据集
- 大模型在混写指令上表现普遍下降,混写密度越高越差
- 推理任务受干扰较小,因生成解释可提供上下文支持
罗马化代码混用(RCM)是多语言社区中以拉丁字母混合本地语言与英语的主流沟通方式。尽管大语言模型(LLMs)在单语和原生脚本任务上表现优异,但其对基于RCM内容的指令理解与推理能力仍缺乏系统评估。为此,我们提出Indi-RomCoM基准,用于评估印地语系罗马化混合指令下的模型表现。该基准涵盖七类指令跟随任务、四种广泛使用的印地语系语言及三种可控代码混用强度。我们在零样本和少样本设置下,全面评测了包括专有模型、开源模型和专注印地语的模型在内的多种LLMs。结果显示,所有模型在RCM指令上的表现均显著下降,且随着代码混用密度增加而持续恶化。值得注意的是,推理任务受损程度低于检测任务(如毒性识别),因生成的解释能提供必要上下文。我们认为Indi-RomCoM有助于推动更具包容性的多语言系统发展。
原文摘要 · Abstract (English)
Romanized Code Mixing (RCM), where bilingual speakers fluidly blend local languages with English in Roman script, has emerged as the dominant form of communication across multilingual communities. While Large Language Models (LLMs) perform strongly on monolingual and native-script benchmarks, their ability to follow instructions and reason over RCM-based content remains largely unexplored. To this end, we introduce the Indi-RomCoM benchmark for facilitating systematic evaluation on Indic Romanized Code-Mixed instructions. Our benchmark spans seven instruction-following tasks, four widely spoken Indic languages, and three controlled code-mixing intensity levels. We extensively evaluate a suite of LLMs covering proprietary, open-weight, and Indic-focused models under zero- and few-shot settings. LLMs consistently underperform on RCM instructions, with performance degrading as code-mixing density increases. Furthermore, reasoning tasks suffer less degradation than detection tasks (e.g., Toxicity) because the generated explanations offer necessary context. We believe Indi-RomCoM helps the community in developing inclusive multilingual systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。