提出新指标衡量上下文对语言模型的说服力。
How Persuasive is Your Context?
- 用沃尔德斯坦距离量化上下文如何改变模型答案分布。
- 实验表明该指标比传统方法更精细捕捉说服效果。
- 适合研究模型行为可塑性与对抗攻击的学者使用。
语言模型的两大核心能力是:(i)调用对实体的先验知识,如回答“奥地利的官方语言是什么?”;(ii)根据上下文中提供的新信息进行调整,例如“假设奥地利的官方语言是塔加洛语”,该信息会前置在问题前。本文引入目标说服度(TPS),用于量化给定上下文对语言模型的说服力,其中说服力被定义为上下文使模型答案发生变化的能力。与仅通过贪婪解码结果评估说服力不同,TPS 提供了更细粒度的模型行为分析。基于沃尔德斯坦距离,TPS 衡量上下文如何将模型原始答案分布推向目标分布。通过一系列实验证明,TPS 能够比以往指标更细致地捕捉说服性的内涵。
原文摘要 · Abstract (English)
Two central capabilities of language models (LMs) are: (i) drawing on prior knowledge about entities, which allows them to answer queries such as "What's the official language of Austria?", and (ii) adapting to new information provided in context, e.g., "Pretend the official language of Austria is Tagalog.", that is pre-pended to the question. In this article, we introduce targeted persuasion score (TPS), designed to quantify how persuasive a given context is to an LM where persuasion is operationalized as the ability of the context to alter the LM's answer to the question. In contrast to evaluating persuasiveness only by inspecting the greedily decoded answer under the model, TPS provides a more fine-grained view of model behavior. Based on the Wasserstein distance, TPS measures how much a context shifts a model's original answer distribution toward a target distribution. Empirically, through a series of experiments, we show that TPS captures a more nuanced notion of persuasiveness than previously proposed metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。