arXiv:2503.13510cs.CLcs.AI2025-03被引 12

研究提示语情感如何影响大模型输出质量,揭示负面提示降低准确率、放大偏见。

Prompt Sentiment: The Catalyst for LLM Change

  • 用词典与Transformer方法分类提示语情感,评估五款主流大模型响应
  • 负面提示使事实准确性下降,偏见加剧;正面提示更冗长且传播情绪
  • 适用于需要公平可靠生成内容的AI应用开发者与提示工程实践者

大语言模型(LLMs)已彻底改变自然语言处理,但输入文本中隐含的情感特征——提示语情感——的影响仍缺乏系统研究。本研究系统考察了提示语情感变化对大模型生成内容在连贯性、事实性和偏见方面的影。我们采用词典和Transformer两种情感分析方法对提示语进行分类,并评估了Claude、DeepSeek、GPT-4、Gemini和LLaMA五款领先模型在六类应用场景中的输出表现:内容生成、对话式AI、法律金融分析、医疗AI、创意写作和技术文档。通过改写提示语,分析其对输出质量的影响。结果表明,提示语情感显著影响模型响应:负面提示常导致事实准确性下降并放大偏见,而正面提示则增加表达冗余与情绪传播。研究强调了在生成可信、公平内容时,进行情感敏感的提示工程的重要性。

原文摘要 · Abstract (English)

The rise of large language models (LLMs) has revolutionized natural language processing (NLP), yet the influence of prompt sentiment, a latent affective characteristic of input text, remains underexplored. This study systematically examines how sentiment variations in prompts affect LLM-generated outputs in terms of coherence, factuality, and bias. Leveraging both lexicon-based and transformer-based sentiment analysis methods, we categorize prompts and evaluate responses from five leading LLMs: Claude, DeepSeek, GPT-4, Gemini, and LLaMA. Our analysis spans six AI-driven applications, including content generation, conversational AI, legal and financial analysis, healthcare AI, creative writing, and technical documentation. By transforming prompts, we assess their impact on output quality. Our findings reveal that prompt sentiment significantly influences model responses, with negative prompts often reducing factual accuracy and amplifying bias, while positive prompts tend to increase verbosity and sentiment propagation. These results highlight the importance of sentiment-aware prompt engineering for ensuring fair and reliable AI-generated content.

大模型提示工程情感分析生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。