arXiv:2410.12631cs.AIcs.LG2024-10被引 2

用知识推理+大模型生成,让道德价值分类结果可解释。

Explainable Moral Values: a neuro-symbolic approach to value classification

  • 基于道德基础理论构建知识图谱,用符号推理识别句子中的价值
  • 仅靠推理已实现可解释分类,融合语义方法效果超越复杂神经网络
  • 适合研究伦理、可解释AI的学者,工具开源可交互

本文探索将基于本体的推理与机器学习结合,实现可解释的价值分类。依托道德基础理论(Moral Foundations Theory)和DnS本体设计模式,使用sandra神经符号推理器,将句子形式化为满足特定价值的描述,并自动推断其对应的价值。句子及其结构化表示通过开源大语言模型自动生成。实验表明,仅依赖推理器的输出即可实现可解释性分类,性能媲美复杂方法;而将推理结果与分布语义方法结合,显著优于所有基线,包括复杂的神经网络模型。此外,作者开发了可视化工具,用于探索基于理论的价值分类,该工具已公开于http://xmv.geomeaning.com/。

原文摘要 · Abstract (English)

This work explores the integration of ontology-based reasoning and Machine Learning techniques for explainable value classification. By relying on an ontological formalization of moral values as in the Moral Foundations Theory, relying on the DnS Ontology Design Pattern, the \textit{sandra} neuro-symbolic reasoner is used to infer values (fomalized as descriptions) that are \emph{satisfied by} a certain sentence. Sentences, alongside their structured representation, are automatically generated using an open-source Large Language Model. The inferred descriptions are used to automatically detect the value associated with a sentence. We show that only relying on the reasoner's inference results in explainable classification comparable to other more complex approaches. We show that combining the reasoner's inferences with distributional semantics methods largely outperforms all the baselines, including complex models based on neural network architectures. Finally, we build a visualization tool to explore the potential of theory-based values classification, which is publicly available at http://xmv.geomeaning.com/.

可解释性知识推理道德分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。