arXiv:2608.29215cs.CL2026-08中稿 · EMNLP

用属性向量精准引导大模型生成符合特定群体的解释。

Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

论文配图:Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation
图 1 · 摘自论文原文
  • 基于目标群体的表达风格与知识水平,计算属性导向的调节向量。
  • 相比提示和现有方法,生成解释更贴合目标群体且事实性保持良好。
  • 适合需要个性化解释的教育、科普等场景使用。

为有效帮助人们理解新主题,解释内容应根据其背景和能力进行定制。目前仅靠提示(prompting)已证明不足以实现这种定制化,且缺乏其他计算方法。因此,本文研究了是否可引导大语言模型(LLM)生成针对特定群体的解释。为此,我们提出一种方法:首先识别目标群体在解释风格和知识水平上的特征属性;基于激活工程,计算属性相关的调节向量,并在推理过程中将其注入模型内部激活中,实现细粒度的引导。实验评估了该方法在解释特异性与事实性方面的有效性。此外,我们还通过不同目标群体的人类专家进行了对照研究。结果表明,相较于提示法和当前最优的引导基线,本方法能更显著地将解释适配至目标群体,同时基本保持事实正确性。

原文摘要 · Abstract (English)

To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. Prompting alone has been shown to be insufficient for creating such explanations, and other computational methods are missing so far. Therefore, this paper investigates whether LLMs can be steered to generate explanations that are tailored to a specific group of people. To this end, we propose an approach that first identifies group-specific attributes in terms of explanatory style and knowledge of a specific target group. Building on activation engineering, it then computes attribute-based steering vectors and adds them to the internal activations of an LLM during inference to enable a fine-grained steering. In our experiments, we assess the steering effectiveness in terms of specificity and factuality of the generated explanations. Additionally, we evaluate the explanations in a study with human experts from different target groups. Compared to prompting and state-of-the-art steering baselines, our approach tailors the explanations significantly better to the target group while maintaining the best specificity-factuality balance.

大模型引导个性化解释群体适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。