arXiv:2601.10205cs.CLcs.AI2026-01

测试12种印地语系语言中嵌入模型对人物与指令的匹配能力。

One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?

  • 构建跨语言人物-指令匹配基准,分离检索与生成任务。
  • 最高召回率27.4%(单语言),32.1%(反向检索)。
  • 为低资源印度语言模型选型提供实证依据。

将多语言助手与文化背景用户偏好对齐,对服务超过十亿使用者、涵盖多种文字的印度语言多样性至关重要。现有基准或仅关注单一语言,或混淆检索与生成,未解答当前嵌入模型能否在不依赖回复生成的前提下编码人物-指令兼容性。本文提出一个涵盖12种印度语言的统一基准,包含四项评估任务:单语与跨语言人物到指令检索、指令到人物反向检索,以及二元兼容性分类。在冻结编码器设置下,使用轻量逻辑回归头评估八种多语言嵌入模型。E5-Large-Instruct 在单语言检索中达到27.4% Recall@1,跨语言迁移达20.7%;BGE-M3 在反向检索中表现最佳,达32.1% Recall@1;LaBSE 在分类任务中取得75.3% AUROC,且校准性能良好。研究结果为印地语系多语言检索模型选择提供实践指导,并建立可复现基线。代码、数据集与模型已公开于 https://github.com/aryashah2k/PI-Indic-Align。

原文摘要 · Abstract (English)

Aligning multilingual assistants with culturally grounded user preferences is essential for serving India's linguistically diverse population of over one billion speakers across multiple scripts. However, existing benchmarks either focus on a single language or conflate retrieval with generation, leaving open the question of whether current embedding models can encode persona-instruction compatibility without relying on response synthesis. We present a unified benchmark spanning 12 Indian languages and four evaluation tasks: monolingual and cross-lingual persona-to-instruction retrieval, reverse retrieval from instruction to persona, and binary compatibility classification. Eight multilingual embedding models are evaluated in a frozen-encoder setting with a thin logistic regression head for classification. E5-Large-Instruct achieves the highest Recall@1 of 27.4\% on monolingual retrieval and 20.7\% on cross-lingual transfer, while BGE-M3 leads reverse retrieval at 32.1\% Recall@1. For classification, LaBSE attains 75.3\% AUROC with strong calibration. These findings offer practical guidance for model selection in Indic multilingual retrieval and establish reproducible baselines for future work\footnote{Code, datasets, and models are publicly available at https://github.com/aryashah2k/PI-Indic-Align.

多语言嵌入对齐低资源印度语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。