arXiv:2605.23069cs.CL2026-05ACL被引 1

通过激活调控提升多语言模型的文化认知能力,无需参数更新。

DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge

论文配图:DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge
图 1 · 摘自论文原文
  • 用平行语料提取语言向量,在推理时注入残差流进行定向调控。
  • 在多选题任务中达到86.96%准确率,排名第七。
  • 揭示了调控效果依赖层位置、语言-地区组合和提示设计。

大型语言模型在多语言与文化场景中应用日益广泛,但其文化知识分布不均。本文提出DFKI-MLT系统参与SemEval-2026任务7(文化意识),通过FLORES平行语料提取语言向量,对多语言LLM实施激活调控。方法在推理阶段于特定Transformer层向残差流添加语言特异性调控向量,无需参数更新。仅多选题(MCQ)提交获得官方评分,准确率达86.96%,位列17支队伍中的第7名。后验分析显示,激活调控对文化推理带来有限且异质性提升:效果高度依赖层位置,跨语言-区域组合差异显著,部分配置甚至导致性能下降,并受提示形式影响,相比通用提示,文化引导提示表现更优。结果表明,提示设计与激活调控需协同优化以实现文化感知的多语言推理。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used across diverse linguistic and cultural contexts, yet their cultural knowledge remains uneven across regions and languages. We present the DFKI-MLT system for SemEval-2026 Task 7 on cultural awareness, where we apply activation steering to multilingual LLMs using language vectors extracted from parallel FLORES data. Our method performs inference-time adaptation by adding language-specific steering vectors to the residual stream at a selected transformer layer, without any parameter updates. We participated in both the short-answer (SAQ) and multiple-choice (MCQ) tracks; however, only our MCQ submission received an official score. In the official MCQ track, we achieved 86.96% accuracy, ranking 7th out of 17 teams. To better understand system behavior, we conduct post-hoc analyses on the shared-task MCQ and SAQ settings. These analyses show that activation steering yields modest and heterogeneous improvements on cultural reasoning: gains are strongly layer-sensitive, vary substantially across language-region pairs, with some configurations even degrading performance, and interact with prompt formulation, comparing generic and culturally conditioned prompts. Our findings suggest that prompt design and activation steering should be jointly optimized for culturally aware multilingual inference.

多语言模型文化认知激活调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。