揭示多语言模型处理印英混用文本时的语义偏倚,提出新方法实现双语言平衡对齐。
Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders
- 通过对比分析发现模型更依赖英语语义空间,导致混用文本与母语间连接松散。
- 在印英混用数据上继续预训练可增强英-混用对齐,但会削弱英-印对齐。
- 提出三语对齐目标,使模型同时贴近两种语言,提升情感分析和仇恨言论检测效果。
多语言编码器广泛用于混用语言分析任务,但我们对其内部如何表征混用输入知之甚少,尤其不清楚这些表征是否真正关联到构成语言。以印地语-英语混用为例,我们构建了一个包含并行英文、天城文印地语和罗马化混用句的统一三语语料库,并通过CKA、词元级显著性及基于熵的不确定性分析,探测标准多语言编码器及其混用适配变体中的跨语言表征对齐情况。结果表明:虽然标准模型能良好对齐英文与印地语,但混用输入仍与任一语言连接松散;在混用数据上持续预训练虽提升了英文-混用对齐,却损害了英文-印地语对齐。可解释性分析进一步揭示明显不对称性:模型通过以英语为主导的语义子空间处理混用文本,而原生脚本印地语则提供补充信号以降低表征不确定性。基于此,我们引入一种三语后训练对齐目标,使混用表征同时更贴近两种构成语言,实现更均衡的跨语言对齐,并在情感分析和仇恨言论检测任务中取得下游性能提升,证明将混用表征锚定于其构成语言能有效促进跨语言理解。
原文摘要 · Abstract (English)
Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally - or whether those representations meaningfully connect to the constituent languages being mixed. Using Hindi-English as a case study, we construct a unified trilingual corpus of parallel English, Hindi (Devanagari), and Romanized code-mixed sentences, and probe cross-lingual representation alignment across standard multilingual encoders and their code-mixed adapted variants via CKA, token-level saliency, and entropy-based uncertainty analysis. We find that while standard models align English and Hindi well, code-mixed inputs remain loosely connected to either language - and that continued pre-training on code-mixed data improves English-code-mixed alignment at the cost of English-Hindi alignment. Interpretability analyses further reveal a clear asymmetry: models process code-mixed text through an English-dominant semantic subspace, while native-script Hindi provides complementary signals that reduce representational uncertainty. Motivated by these findings, we introduce a trilingual post-training alignment objective that brings code-mixed representations closer to both constituent languages simultaneously, yielding more balanced cross-lingual alignment and downstream gains on sentiment analysis and hate speech detection - showing that grounding code-mixed representations in their constituent languages meaningfully helps cross-lingual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。