探索提示词具体性对专业领域模型表现的影响,发现存在最佳词汇精确度范围。
Prompt Engineering: How Prompt Vocabulary affects Domain Knowledge
- 通过替换名词动词形容词的精确度,系统测试提示词具体性影响
- 多数模型在适度具体提示下表现最佳,过度具体反而降低效果
- 适合需要高效精准问答的科研、医疗、法律等专业场景应用
提示工程已成为优化大语言模型(LLMs)在特定领域任务中表现的关键因素。然而,在科学、技术、工程、数学(STEM)、医学和法律等领域的提示词具体性作用仍研究不足。本论文探究增加提示词词汇具体性是否能提升LLM在领域问答与推理任务中的表现。我们构建了同义词替换框架,系统性地将名词、动词和形容词替换为不同具体程度的词汇,评估四种模型(Llama-3.1-70B-Instruct、Granite-13B-Instruct-V2、Flan-T5-XL、Mistral-Large 2)在STEM、法律和医学数据集上的表现。结果表明,虽然普遍提高提示词具体性无显著提升,但在所有模型中均存在一个最佳具体性范围,此时模型表现最优。该发现为提示设计提供关键洞见:在该范围内调整提示可最大化模型性能,推动其在专业领域的高效应用。
原文摘要 · Abstract (English)
Prompt engineering has emerged as a critical component in optimizing large language models (LLMs) for domain-specific tasks. However, the role of prompt specificity, especially in domains like STEM (physics, chemistry, biology, computer science and mathematics), medicine, and law, remains underexplored. This thesis addresses the problem of whether increasing the specificity of vocabulary in prompts improves LLM performance in domain-specific question-answering and reasoning tasks. We developed a synonymization framework to systematically substitute nouns, verbs, and adjectives with varying specificity levels, measuring the impact on four LLMs: Llama-3.1-70B-Instruct, Granite-13B-Instruct-V2, Flan-T5-XL, and Mistral-Large 2, across datasets in STEM, law, and medicine. Our results reveal that while generally increasing the specificity of prompts does not have a significant impact, there appears to be a specificity range, across all considered models, where the LLM performs the best. Identifying this optimal specificity range offers a key insight for prompt design, suggesting that manipulating prompts within this range could maximize LLM performance and lead to more efficient applications in specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。