Ethic-BERT提升伦理内容分类准确率,更懂复杂道德判断。
Ethic-BERT: An Enhanced Deep Learning Model for Ethical and Non-Ethical Content Classification
- 基于BERT改进,融合预处理与动态调优策略
- 标准测试准确率达82.32%,难题集提升15.28%
- 适合需要可靠道德推理的AI系统开发
开发具备细致伦理推理能力的AI系统至关重要,因它们日益影响人类决策,但现有模型常依赖表面关联而非原则性道德理解。本文提出Ethic-BERT,一种基于BERT的模型,用于跨四个领域(常识、公正、美德、义务论)的伦理与非伦理内容分类。利用ETHICS数据集,该方法结合强预处理以缓解词汇稀疏与语境模糊,并采用全模型解冻、梯度累积和自适应学习率调度等先进微调策略。为评估鲁棒性,引入对抗过滤的“难例测试”划分,分离出复杂伦理困境。实验表明,Ethic-BERT在标准测试中平均准确率达82.32%,在公正与美德领域表现尤为突出;在难例测试中,平均准确率提升15.28%。研究结果推动了性能提升与基于偏差感知预处理的可靠决策。
原文摘要 · Abstract (English)
Developing AI systems capable of nuanced ethical reasoning is critical as they increasingly influence human decisions, yet existing models often rely on superficial correlations rather than principled moral understanding. This paper introduces Ethic-BERT, a BERT-based model for ethical content classification across four domains: Commonsense, Justice, Virtue, and Deontology. Leveraging the ETHICS dataset, our approach integrates robust preprocessing to address vocabulary sparsity and contextual ambiguities, alongside advanced fine-tuning strategies like full model unfreezing, gradient accumulation, and adaptive learning rate scheduling. To evaluate robustness, we employ an adversarially filtered "Hard Test" split, isolating complex ethical dilemmas. Experimental results demonstrate Ethic-BERT's superiority over baseline models, achieving 82.32% average accuracy on the standard test, with notable improvements in Justice and Virtue. In addition, the proposed Ethic-BERT attains 15.28% average accuracy improvement in the HardTest. These findings contribute to performance improvement and reliable decision-making using bias-aware preprocessing and proposed enhanced AI model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。