用显式推理视角引导大模型,提升不当言论检测准确率
Soft Inductive Bias Approach via Explicit Reasoning Perspectives in Inappropriate Utterance Detection Using Large Language Models
- 通过定义推理视角构建软归纳偏置,约束模型推理过程
- 在韩语数据上实现87.0046%平均准确率,较传统方法提升3.89%
- 适合需要可解释性与稳定判断的在线内容安全场景
某些保障匿名性的在线游戏和社区中,未经管控的不当言论频繁升级为言语攻击甚至犯罪行为,引发重大社会关切。因此亟需研究对话文本中的不当言论检测技术,以构建更安全的交流环境。尽管基于韩语语料训练的大规模语言模型及思维链推理近期受到关注,但将其应用于不当言论检测的研究仍有限。本文提出一种软归纳偏置方法,通过显式定义推理视角来引导推理过程,促进理性决策并避免推理错误。我们采用该方法微调韩语大模型,并在不同训练策略下进行定量性能对比与定性评估。实验结果表明,Kanana-1.5模型平均准确率达87.0046%,相比标准监督学习提升约3.89%。研究显示,该方法超越了大模型对知识的简单模仿,通过受限的推理视角实现更精准、一致的判断,验证了其在不当言论检测中的有效性。
原文摘要 · Abstract (English)
Recent incidents in certain online games and communities, where anonymity is guaranteed, show that unchecked inappropriate remarks frequently escalate into verbal abuse and even criminal behavior, raising significant social concerns. Consequently, there is a growing need for research on techniques that can detect inappropriate utterances within conversational texts to help build a safer communication environment. Although large-scale language models trained on Korean corpora and chain-of-thought reasoning have recently gained attention, research applying these approaches to inappropriate utterance detection remains limited. In this study, we propose a soft inductive bias approach that explicitly defines reasoning perspectives to guide the inference process, thereby promoting rational decision-making and preventing errors that may arise during reasoning. We fine-tune a Korean large language model using the proposed method and conduct both quantitative performance comparisons and qualitative evaluations across different training strategies. Experimental results show that the Kanana-1.5 model achieves an average accuracy of 87.0046, improving by approximately 3.89 percent over standard supervised learning. These findings indicate that the proposed method goes beyond simple knowledge imitation by large language models and enables more precise and consistent judgments through constrained reasoning perspectives, demonstrating its effectiveness for inappropriate utterance detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。