发现大模型难以同时控制幽默与说服力等概念
Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs
- 设计新评估框架,测试单/双概念精细控制能力
- 多模型实验显示双概念控制性能普遍下降
- 揭示提示工程在组合控制上的根本缺陷
大型语言模型具备强大的生成能力,但许多应用需要对特定文本概念(如幽默、说服力或正式性)进行明确且细粒度的控制。现有提示工程与表征设计方法仅能实现粗粒度或单一属性控制,对多属性设置的系统性评估仍不足。本文提出一种针对单概念和双概念场景的细粒度可控性评估框架,聚焦语言上差异明显的概念对(如说服力与幽默)。实验发现,尽管这些概念理论上可分离,但在多个大模型和生成任务中,双概念控制下的性能普遍下降。这一现象揭示了基于提示的控制存在根本局限:即使概念直观独立,模型仍难以处理组合性。该框架为未来多概念控制方法的能力测量提供了系统性证据和原则性路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer strong generative capabilities, but many applications require explicit and \textit{fine-grained} control over specific textual concepts, such as humor, persuasiveness, or formality. Prior approaches in prompting and representation engineering can provide coarse or single-attribute control, but systematic evaluation of multi-attribute settings remains limited. We introduce an evaluation framework for fine-grained controllability for both single- and dual-concept scenarios, focusing on linguistically distinct concept pairs (e.g., persuasiveness vs.~humor). Surprisingly, across multiple LLMs and generative tasks, we find that performance often drops in the dual-concept setting, even though the chosen concepts should in principle be separable. This reveals a fundamental limitation of naive prompting-based control: models struggle with compositionality even when concepts are intuitively independent. Our framework provides systematic evidence of this gap and offers a principled approach for measuring the ability of future methods for multi-concept control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。