arXiv:2603.02578cs.CLcs.AI2026-03ACL被引 3

提出分层评估框架,检验大模型在语言、情感和人格上的可控性。

How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities

  • 构建三级控制层级:表达什么、如何表达、如何实现。
  • 发现细粒度控制下模型表现明显下降。
  • 适合关注安全可控的AI研发人员使用。

大语言模型日益应用于敏感社会领域,但其行为不可预测,从意图错位到人格不一致均带来风险。本文提出SteerEval,一个分层基准,用于评估模型在语言特征、情感和人格三个领域的可控性。每个领域设三级规范:L1(表达内容)、L2(表达方式)、L3(具体实现),将高层意图与具体文本输出相连接。通过该基准系统评估现有控制方法,发现控制能力在细粒度层级显著退化。该框架为安全可控的模型行为提供可解释的评估基础,推动后续研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in socially sensitive domains, yet their unpredictable behaviors, ranging from misaligned intent to inconsistent personality, pose significant risks. We introduce SteerEval, a hierarchical benchmark for evaluating LLM controllability across three domains: language features, sentiment, and personality. Each domain is structured into three specification levels: L1 (what to express), L2 (how to express), and L3 (how to instantiate), connecting high-level behavioral intent to concrete textual output. Using SteerEval, we systematically evaluate contemporary steering methods, revealing that control often degrades at finer-grained levels. Our benchmark offers a principled and interpretable framework for safe and controllable LLM behavior, serving as a foundation for future research.

大模型可控性评估基准行为控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。