arXiv:2510.12728cs.HCcs.LG2025-10

让测试数据和提示词一起进化,更精准地调整大模型行为。

Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior

  • 测试集与提示词同步迭代,形成动态反馈闭环。
  • 用户研究显示该方法能系统化提升提示词对齐度。
  • 适合需要精细控制模型行为的应用开发者使用。

大型语言模型在应用中日益普及,用户可通过修改提示词来调整模型行为。然而,将细微的领域特定策略编码进提示词颇具挑战性。尽管这一过程通常需要具体的测试案例,但测试数据与提示词往往作为独立产物开发,反映了传统机器学习中模型调优缓慢、测试集静态的实践。我们主张,提示工程快速迭代的特性要求打破这种分离,推动一种新工作流:数据-提示协同演化,即动态测试集与提示词共同演进。本文提出一个交互式系统,引导开发者发现边缘案例、阐明期望行为理由,并针对不断增长的测试集迭代评估修订后的提示词。用户研究表明,该工作流有助于人们系统性地优化提示词,使其更符合预期政策。本工作通过人机协作开发,指向更稳健、更负责任的大模型应用。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly embedded in applications, and people can shape model behavior by editing prompt instructions. Yet encoding subtle, domain-specific policies into prompts is challenging. Although this process often benefits from concrete test cases, test data and prompt instructions are typically developed as separate artifacts, reflecting traditional machine learning practices in which model tuning was slow and test sets were static. We argue that the fast, iterative nature of prompt engineering calls for removing this separation and enabling a new workflow: data-prompt co-evolution, where a living test set and prompt instructions evolve in tandem. We present an interactive system that operationalizes this workflow. It guides application developers to discover edge cases, articulate rationales for desired behavior, and iteratively evaluate revised prompts against a growing test set. A user study shows our workflow helps people refine prompts systematically, better aligning them with their intended policies. This work points toward more robust and responsible LLM applications through human-in-the-loop development.

提示工程人机协作模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。