评测大模型通过提示词切换人格的能力,发现多数模型受限于初始行为偏移和调节不对称。
Evaluating the Prompt Steerability of Large Language Models
- 基于提示词对模型行为分布的可调性,定义量化评估指标
- 多数模型在多个人格维度上存在调节能力不足现象
- 适合关注模型可控性与价值观对齐的研究者
构建多元化的人工智能需要设计能够体现多种价值体系和文化背景的模型。实现这一目标的前提是评估模型通过提示词调节其人格表现的能力。为此,我们提出一个基准测试方法,用于评估模型人格在提示词驱动下的可调节性。该方法基于提示词可调性的形式化定义,分析模型联合行为分布从基线状态被改变的程度。通过定义可调性指数,并观察这些指数随调节努力的变化,我们可以估计模型在不同人格维度和方向上的可调性。实验结果表明,许多现有模型的可调性受限——既由于其基线行为本身的偏差,也由于在多个性格维度上调节能力的非对称性。相关代码已开源:https://github.com/IBM/prompt-steering。
原文摘要 · Abstract (English)
Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model is capable of reflecting various personas. To this end, we propose a benchmark for evaluating the steerability of model personas as a function of prompting. Our design is based on a formal definition of prompt steerability, which analyzes the degree to which a model's joint behavioral distribution can be shifted from its baseline. By defining steerability indices and inspecting how these indices change as a function of steering effort, we can estimate the steerability of a model across various persona dimensions and directions. Our benchmark reveals that the steerability of many current models is limited -- due to both a skew in their baseline behavior and an asymmetry in their steerability across many persona dimensions. We release an implementation of our benchmark at https://github.com/IBM/prompt-steering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。