arXiv:2604.04869cs.LG2026-04

用声明式框架自动优化大模型提示词,提升准确率与可靠性。

Optimizing LLM Prompt Engineering with DSPy Based Declarative Learning

  • 通过声明式框架自动构建和优化提示词模块。
  • 事实准确率提升30%~45%,幻觉率降低约25%。
  • 适合需要高可靠性的复杂推理任务研究者。

大型语言模型在自然语言处理任务中表现出色,但其效果高度依赖于提示词设计、结构及内嵌的推理信号。传统提示工程方法主要依赖经验性试错,限制了可扩展性、可复现性以及跨任务泛化能力。本文提出基于DSPy的声明式学习系统,实现提示词的自动化、模块化与可学习构建。我们引入统一的DSPy LLM架构,结合符号规划、无梯度优化与模块自动重写机制,有效减少幻觉、提升事实准确性并避免提示冗余。在推理任务、检索增强生成及多步思维链基准测试中,实验表明该方法在输出可靠性、效率和泛化能力方面均有显著提升,事实准确率最高提升30%至45%,幻觉率下降约25%。最后,我们讨论了当前局限与未来研究方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown strong performance across a wide range of natural language processing tasks; however, their effectiveness is highly dependent on prompt design, structure, and embedded reasoning signals. Conventional prompt engineering methods largely rely on heuristic trial-and-error processes, which limits scalability, reproducibility, and generalization across tasks. DSPy, a declarative framework for optimizing text-processing pipelines, offers an alternative approach by enabling automated, modular, and learnable prompt construction for LLM-based systems.This paper presents a systematic study of DSPy-based declarative learning for prompt optimization, with emphasis on prompt synthesis, correction, calibration, and adaptive reasoning control. We introduce a unified DSPy LLM architecture that combines symbolic planning, gradient free optimization, and automated module rewriting to reduce hallucinations, improve factual grounding, and avoid unnecessary prompt complexity. Experimental evaluations conducted on reasoning tasks, retrieval-augmented generation, and multi-step chain-of-thought benchmarks demonstrate consistent gains in output reliability, efficiency, and generalization across models. The results show improvements of up to 30 to 45% in factual accuracy and a reduction of approximately 25% in hallucination rates. Finally, we outline key limitations and discuss future research directions for declarative prompt optimization frameworks.

提示工程大模型声明式学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。