arXiv:2510.24236cs.CL2025-10被引 5

研究如何让大模型解释更可信,尤其在医疗场景中。

Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?

  • 通过调整提示词和少量示例来提升模型解释的可靠性
  • 少样本示例的数量与质量显著影响解释准确性
  • 指令微调能有效提高医疗问答任务中的可信度

大语言模型常生成不反映其预测依据的解释,在医疗场景中尤为危险:遗漏关键临床线索或依赖虚假捷径会削弱医生信任并导致决策失误。本文研究推理与训练阶段的选择如何影响解释的忠实性,聚焦部署时可控制的因素。在BBQ(社会偏见)和MedQA(医学执照考试)两个数据集上,测试GPT-4.1-mini、LLaMA 70B和LLaMA 8B三种模型,操纵少量示例数量与类型、提示策略及训练流程。结果表明:(i) 少样本示例的数量与质量对模型忠实性有显著影响;(ii) 提示设计敏感影响解释可信度;(iii) 指令微调阶段提升了MedQA任务上的测量忠实性。研究为提升敏感领域中模型可解释性与可信度提供了实用策略。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often produce explanations that do not faithfully reflect the factors driving their predictions. In healthcare settings, such unfaithfulness is especially problematic: explanations that omit salient clinical cues or mask spurious shortcuts can undermine clinician trust and lead to unsafe decision support. We study how inference and training-time choices shape explanation faithfulness, focusing on factors practitioners can control at deployment. We evaluate three LLMs (GPT-4.1-mini, LLaMA 70B, LLaMA 8B) on two datasets-BBQ (social bias) and MedQA (medical licensing questions), and manipulate the number and type of few-shot examples, prompting strategies, and training procedure. Our results show: (i) both the quantity and quality of few-shot examples significantly impact model faithfulness; (ii) faithfulness is sensitive to prompting design; (iii) the instruction-tuning phase improves measured faithfulness on MedQA. These findings offer insights into strategies for enhancing the interpretability and trustworthiness of LLMs in sensitive domains.

大模型解释医疗AI可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。