用量化指标和迭代提示提升LLM写研究计划书的质量与可信度
Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement
- 设计内容质量与参考文献有效性双指标,实现客观评估
- 迭代提示使内容质量提升,虚假引用减少70%以上
- 适合科研人员评估和改进AI辅助写作成果
大型语言模型(如ChatGPT)在学术写作中应用日益广泛,但存在引用错误或虚构参考文献等伦理问题。现有内容质量评估多依赖主观人工判断,耗时且缺乏客观性。本研究提出内容质量与参考文献有效性两项量化评估指标,并基于评分设计迭代提示方法。大量实验表明,该框架能客观衡量ChatGPT的写作表现;迭代提示显著提升内容质量,同时将参考文献错误与虚构率降低70%以上,有效缓解学术场景中的关键伦理挑战。
原文摘要 · Abstract (English)
Large language models (LLMs) like ChatGPT are increasingly used in academic writing, yet issues such as incorrect or fabricated references raise ethical concerns. Moreover, current content quality evaluations often rely on subjective human judgment, which is labor-intensive and lacks objectivity, potentially compromising the consistency and reliability. In this study, to provide a quantitative evaluation and enhance research proposal writing capabilities of LLMs, we propose two key evaluation metrics--content quality and reference validity--and an iterative prompting method based on the scores derived from these two metrics. Our extensive experiments show that the proposed metrics provide an objective, quantitative framework for assessing ChatGPT's writing performance. Additionally, iterative prompting significantly enhances content quality while reducing reference inaccuracies and fabrications, addressing critical ethical challenges in academic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。