人在无标准答案时优化提示词效果差,多数人迭代后反而更糟。
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
- 通过表格工具让用户反复优化提示词标注数据
- 20人中仅9人经4轮以上迭代后准确率提升
- 无标准标签时自动化工具也难奏效,需谨慎依赖
数百万用户使用大语言模型(LLMs)完成各类任务,但人类在提示工程上的表现如何?在缺乏金标准标签的情况下,用户能否通过多次迭代提示词逼近目标结果?本研究考察了「无光提示」场景——即用户在无手工标注基准的情况下,通过迭代提示词进行数据标注。我们开发了 PromptingSheet,一款 Google Sheets 插件,支持用户在表格中编写、修改并逐步标注数据。对20名参与者的实验发现,在无金标的情况下,提示优化极不可靠:仅9名参与者在四轮及以上迭代后提升了标注准确率。即使采用 DSPy 等自动化提示优化工具,在金标极少时仍表现不佳。研究强调了金标准标签的重要性,揭示了人工提示工程中自动化支持的必要性与风险,为未来工具设计提供了关键洞见。
原文摘要 · Abstract (English)
Millions of users prompt large language models (LLMs) for various tasks, but how good are people at prompt engineering? Do users actually get closer to their desired outcome over multiple iterations of their prompts? These questions are crucial when no gold-standard labels are available to measure progress. This paper investigates a scenario in LLM-powered data labeling, "prompting in the dark," where users iteratively prompt LLMs to label data without using manually-labeled benchmarks. We developed PromptingSheet, a Google Sheets add-on that enables users to compose, revise, and iteratively label data through spreadsheets. Through a study with 20 participants, we found that prompting in the dark was highly unreliable -- only 9 participants improved labeling accuracy after four or more iterations. Automated prompt optimization tools like DSPy also struggled when few gold labels were available. Our findings highlight the importance of gold labels and the needs, as well as the risks, of automated support in human prompt engineering, providing insights for future tool design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。