用细粒度评估提升大模型对敏感话题的回答质量
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
- 将回答问题的优劣分解为内容、逻辑、得体性三类错误
- 通过评分反馈使得体性错误率降低33.09%
- 适合希望兼顾安全与实用性的模型优化研究者
大型语言模型在敏感话题上常回应过于谨慎和模糊,牺牲了帮助性以换取安全性。现有评估框架缺乏系统方法来识别和解决响应中的具体缺陷,难以同时提升安全性和帮助性。为此,我们提出FINEST——一种针对敏感话题的细粒度响应评估分类法,将帮助性与无害性拆解为内容、逻辑、得体性三类错误。在韩国敏感问题数据集上的实验表明,基于FINEST的评分与错误导向改进流程显著提升了模型在三类维度的表现,优于无引导的修正方法。尤其,提供类别化评分与理由的评分式改进带来最大提升,使得体性错误句子比例最多下降33.09%。该工作为更可解释、全面的敏感话题响应评估与改进奠定了基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often generate overly cautious and vague responses on sensitive topics, sacrificing helpfulness for safety. Existing evaluation frameworks lack systematic methods to identify and address specific weaknesses in responses to sensitive topics, making it difficult to improve both safety and helpfulness simultaneously. To address this, we introduce FINEST, a FINE-grained response evaluation taxonomy for Sensitive Topics, which breaks down helpfulness and harmlessness into errors across three main categories: Content, Logic, and Appropriateness. Experiments on a Korean-sensitive question dataset demonstrate that our score- and error-based improvement pipeline, guided by FINEST, significantly improves the model responses across all three categories, outperforming refinement without guidance. Notably, score-based improvement -- providing category-specific scores and justifications -- yields the most significant gains, reducing the error sentence ratio for Appropriateness by up to 33.09%. This work lays the foundation for a more explainable and comprehensive evaluation and improvement of LLM responses to sensitive questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。