让解释模型根据听众偏好动态调整用词,提升可懂性。
Direct Preference Optimization for Adaptive Concept-based Explanations
- 基于直接偏好优化,迭代训练说话者模型适应听众
- 在三个数据集上成功对齐模拟听众的偏好,提升理解度
- 适合需要个性化解释的场景,如医疗、教育等
概念化解释方法旨在通过识别输入中对预测任务最重要的语义特征(如颜色、图案、形状)来增强机器学习模型的透明性。然而,这些方法通常忽略解释的沟通语境,例如听者的偏好。例如,医生理解诊断基于临床指标,但患者可能无法理解,需使用不同词汇进行解释。本文提出一种基于语用推理和理性言语行为原则的听众自适应解释方法。通过基于直接偏好优化的迭代训练,使说话者模型生成对听者具有最大沟通效用的解释。该方法仅需成对偏好数据,可由人类反馈获取,在缺乏听众模型的现实场景中尤为适用。我们在三个图像分类数据集上验证了方法能有效对齐模拟听众的偏好,并在用户研究中证明,使用该方法生成的语用解释可显著提升参与者对分类任务的准确率。
原文摘要 · Abstract (English)
Concept-based explanation methods aim at making machine learning models more transparent by finding the most important semantic features of an input (e.g., colors, patterns, shapes) for a given prediction task. However, these methods generally ignore the communicative context of explanations, such as the preferences of a listener. For example, medical doctors understand explanations in terms of clinical markers, but patients may not, needing a different vocabulary to rationalize the same diagnosis. We address this gap with listener-adaptive explanations grounded in principles of pragmatic reasoning and the rational speech act. We introduce an iterative training procedure based on direct preference optimization where a speaker learns to compose explanations that maximize communicative utility for a listener. Our approach only needs access to pairwise preferences, which can be collected from human feedback, making it particularly relevant in real-world scenarios where a model of the listener may not be available. We demonstrate that our method is able to align speakers with the preferences of simulated listeners on image classification across three datasets, and further validate that pragmatic explanations generated with our method improve the classification accuracy of participants in a user study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。