剖析NLP可解释性在七大领域的实际应用与挑战
Explainability in Practice: A Survey of Explainable NLP Across Various Domains
- 按医疗、金融等7大领域梳理解释需求与方法
- 发现不同领域对解释的准确性与成本要求差异显著
- 提出双层评估框架,兼顾技术指标与领域实情
自然语言处理已广泛应用于医疗、金融、客户关系管理等关键领域,GPT-4o、Gemini、BERT等模型日益影响决策。由于这些模型具有黑箱特性,透明性需求迫切。本文系统考察可解释NLP(XNLP)在实际部署中的表现,覆盖医学、金融、系统综述、客户关系管理、聊天机器人、社会行为科学及人力资源七个领域。针对每个领域,分析其解释需求、所用方法及其评估方式,并通过跨领域对比揭示需求差异。比较主流解释方法族在作用范围、忠实度证据和计算成本上的表现。提出两层评估协议:将通用技术指标与领域特定验证层分离,以增强评估有效性。文章还指出当前研究中被忽视的问题,如真实场景适用性、保真度与可信度之间的差距,以及人类判断在解释评估中的作用。最后提出未来方向,包括个性化解释、人机协同评估及大模型机制可解释性。
原文摘要 · Abstract (English)
Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o, Gemini, and BERT increasingly inform decisions. The black-box nature of these models has created an urgent need for transparency. This review examines explainable NLP (XNLP) as it is actually deployed, working through seven application domains: medicine, finance, systematic reviews, customer relationship management, chatbots, social and behavioral science, and human resources. For each domain, we ask what kind of explanation the setting needs, which methods are used there, and how they are evaluated. A structured cross-domain synthesis then contrasts how those requirements diverge. We compare the main explanation method families on scope, evidence of faithfulness, and computational cost. We also propose a two-tier evaluation protocol that separates a shared technical core of metrics from the domain-specific validation layer through which those metrics have to be read. The review also addresses areas that remain underrepresented in the XNLP literature, including real-world applicability, the gap between fidelity and faithfulness, and the role of human judgment in assessing explanations. It closes with research directions, among them personalized explanations, human-in-the-loop evaluation, and mechanistic interpretability for large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。