用符号框架约束大模型,让自动判别仇恨言论更准且可审计
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech

- 将大模型嵌入确定性逻辑树结构,实现法律推理的可解释性
- 在德国刑法第130条下召回率达82%-89%,准确率80%-86%(原方法仅34%-49%)
- 适合需要合规、可审计的在线内容审核场景
自动化法律推理面临两难:符号系统透明但难处理模糊性,神经网络灵活却不可验证。本文探讨混合式神经符号方法能否调和这一矛盾。以在线内容审核为代理场景,模拟大规模行政裁决中每日需处理数千案件的高要求。研究聚焦于将大语言模型(LLMs)置于确定性符号框架内,是否能提升基于法条的非法性判断能力,同时避免“范围漂移”(即混淆道德冒犯与法律违法)。评估了规则映射(Rulemapping)——一种可视化逻辑树方法,用于实现经典法律三段论——在德国刑法§130(1)仇恨言论分类任务上的表现。结果显示,无论何种LLM,Rulemapping均保持0.82-0.89的召回率,精度达0.80-0.86;而未加约束的提示方式精度仅为0.34-0.49。专家构建的符号框架使法律自动化既符合监管审计要求,又具备可验证决策能力。
原文摘要 · Abstract (English)
Automating legal reasoning forces a choice between imperfect alternatives: symbolic systems offer transparency but struggle with ambiguity, whereas neural systems handle natural language flexibly but lack verifiability. This paper investigates whether a hybrid, neuro-symbolic approach can reconcile this trade-off. We evaluate this architecture in the domain of online content moderation, which serves as a proxy for high-volume legal decision-making such as mass administrative proceedings. In these settings, operators must assess thousands of cases daily under strict legal standards. Specifically, we examine whether constraining large language models (LLMs) within deterministic symbolic scaffolds improves statute-grounded illegality assessment while preventing "scope drift" (where LLMs conflate moral offensiveness with legal illegality). We evaluate the neuro-symbolic variant of Rulemapping - a visual logic-tree method that operationalises the classic legal syllogism - on online hate-speech classification under §130(1) of the German Criminal Code. Across diverse LLMs, Rulemapping maintains high recall (0.82-0.89) while achieving precision of 0.80-0.86, compared to 0.34-0.49 for unconstrained prompting. Expert-authored symbolic scaffolds thus enable robust legal automation aligned with regulatory requirements for auditability and verifiable decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。