arXiv:2510.15436cs.CL2025-10被引 14

通过提示工程实现摘要抽象程度可控,提升大模型生成质量。

Controllable Abstraction in Summary Generation for Large Language Models via Prompt Engineering

  • 设计多阶段提示框架,结合语义分析与话题建模控制摘要抽象度。
  • 提示长度适中时效果最佳,过短或过长均降低摘要质量。
  • 对新闻文本效果好,学术文本和噪声数据会显著影响生成结果。

本研究提出一种基于提示工程的可控抽象摘要生成方法,以解决传统方法在摘要质量和可控性方面的不足。设计多阶段提示生成框架,通过对输入文本进行语义分析、话题建模和噪声控制,生成不同抽象层次的摘要。实验采用CNN/Daily Mail数据集,分析提示长度、数据噪声和文本类型的影响。结果表明,提示长度对生成质量有显著影响:过短或过长均导致性能下降。数据噪声随水平升高,ROUGE-L得分逐渐降低。不同文本类型影响模型表现,处理新闻文本时最优,学术文章则较差。研究揭示了优化提示策略与文本预处理对提升摘要准确性和可控性的关键作用。

原文摘要 · Abstract (English)

This study presents a controllable abstract summary generation method for large language models based on prompt engineering. To address the issues of summary quality and controllability in traditional methods, we design a multi-stage prompt generation framework. This framework generates summaries with varying levels of abstraction by performing semantic analysis, topic modeling, and noise control on the input text. The experiment uses the CNN/Daily Mail dataset and provides a detailed analysis of different prompt lengths, data noise, and text types. The experimental results show that prompt length has a significant impact on the quality of generated summaries. Both very short and very long prompt tokens result in a decrease in summary quality. Data noise also negatively affects the summary generation process. As noise levels increase, the ROUGE-L score gradually decreases. Furthermore, different text types have varying effects on the model's ability to generate summaries. The model performs best when handling news texts, while its performance is worse when processing academic articles. This research provides new insights into improving summary generation using large language models, particularly in how controlling prompt strategies and optimizing text preprocessing can enhance summary accuracy and controllability.

摘要生成提示工程可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。