arXiv:2608.14896cs.CLcs.LG2026-08被引 1

探究小模型中日英跨语言语用理解,揭示文化敏感信息的内部表征位置。

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

  • 构建最小差异测试集,分离语用现象与语言流畅性
  • 发现敬语特征在第15层残留流中清晰可辨,准确率达96%
  • 提出无参数推理时编辑方法,可定向调整语用表达

大语言模型在英语表现良好,但在与之类型差异大的语言如日语上行为仍不清晰。现有评估依赖翻译质量与JGLUE类基准,将词汇、句法和语用能力混为一谈。通用模型在日语中的失败多源于语用问题:敬语体系、群体内外指称、情境敏感礼貌、零回指。本文提出J-PragEval-v0,一个隔离四种语用现象的最小对测试集,结合线性探测与教师强制对数概率评估,研究TinySwallow-1.5B(28层,隐藏维度1536)中对应对比的内部位置。四类特征分三类呈现:敬语注册位于残留流中,第15层平衡准确率达0.96,93%样本随场景切换偏好;隐含主语与群内指称在最终提示词处无法线性解码(0.48与0.38),但翻转率分别达0.77与0.79,表明其在生成过程中被处理而非存储于提示端;间接拒绝为反例:探测准确率0.95,但在长度归一化教师强制下降至0.43,因当前最小对同时混入礼貌与延续长度。此外提出语用表征引导(Pragmatic Representation Steering),一种无需参数的推理时编辑方法,沿探测识别的类别均值差方向调整残留流激活。可行性通过基线验证:相同几何结构的对比激活添加,可在存在线性信号处恢复探测准确率至逻辑回归水平±1~2点。扩展至Llama-3.1-Swallow-8B是下一步。

原文摘要 · Abstract (English)

Large language models work well on English and behave in poorly understood ways on languages typologically far from it. Japanese is a clean example, where evaluation still leans on translation quality and JGLUE-style benchmarks, which roll lexical, syntactic and pragmatic competence into a single score. The phenomena on which general-purpose models fail Japanese users are pragmatic: honorifics, in-group and out-group reference, context-sensitive politeness, zero anaphora. I introduce J-PragEval-v0, a minimal-pair benchmark isolating four such phenomena from surface fluency, and combine it with linear probes and teacher-forced log-probability evaluation to ask where inside TinySwallow-1.5B (28 layers, hidden size 1536) the corresponding contrasts live. The four features split three ways. Honorific register sits cleanly in the residual stream: 0.96 balanced accuracy at layer 15, and the model flips its preferred continuation with the scenario on 93 percent of items. Implicit subject and in-group reference are not linearly decodable at the final prompt token (0.48 and 0.38), yet flip rates are 0.77 and 0.79, so the contrast is worked out during generation rather than stored at the prompt. Indirect refusal is the negative case: 0.95 probe accuracy collapsing to a 0.43 flip rate under length-normalised teacher forcing, because the current minimal pairs conflate politeness with continuation length. I also specify Pragmatic Representation Steering, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies. Feasibility is argued indirectly rather than demonstrated: the contrastive activation addition baseline, the same geometry the method would inject, recovers probe accuracy within one to two points of logistic regression wherever a linear signal exists. Scaling to Llama-3.1-Swallow-8B is the next step.

语用分析小模型跨语言可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。