arXiv:2608.25028cs.CLcs.LG2026-08

用提示微调提升痴呆检测准确率,但解释不可信。

Behind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection

  • 将痴呆诊断转为掩码词预测任务,提升模型性能
  • 准确率和宏平均F1达0.83,[MASK]表征最能还原诊断信息
  • 生成的解释主要反映词汇和语篇特征,不具可信度

基于提示的领域自适应模型在低资源、非侵入性痴呆筛查中前景广阔,但其内部机制仍不透明。本文研究了通过提示微调(DAPF)框架实现的领域自适应模型在痴呆检测中的可解释性。该方法将痴呆检测转化为与诊断相关的掩码词预测任务。通过多种探针与分析技术,发现DAPF在整体性能上表现最佳(准确率=0.83,宏平均F1=0.83),且诊断信息最能从其[MASK]表示中恢复。然而,该表示的优势并未延伸至词级解释的可信度。DAPF生成的归因主要反映语言任务词汇、话语标记和转录伪影,扰动测试显示其影响微弱或为负。这表明,尽管掩码接口承载诊断信息,但无法生成可信的词级解释。

原文摘要 · Abstract (English)

Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque. We study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related masked-token prediction. We interpret DAPF and strong baselines using a variety of probing and analysis techniques, finding that DAPF achieved the best overall performance (accuracy=0.83 and macro-F1=0.83) with diagnosis most recoverable from its [MASK] representation. However, this representational advantage did not extend to token-level explanation faithfulness. DAPF attributions primarily reflected language task vocabulary, discourse markers, and transcription artifacts, with perturbation tests showing weak or negative effects. This suggests that its masked-token interface determines diagnosis information without producing faithful token-level explanations.

痴呆检测可解释性提示微调语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。