arXiv:2604.10502cs.AI2026-04ACL

用类比例子提升大模型内容审核的准确性与可解释性

CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMs

论文配图:CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMs
图 1 · 摘自论文原文
  • 通过类比检索与规则生成联合优化,动态生成审核规则
  • 在多个数据集上准确率显著高于基线方法,规则更清晰可靠
  • 适合需要可解释审核策略的平台或监管场景

在线平台的内容审核面临用户生成内容日益复杂、传统规则和机器学习方法局限的挑战。尽管大语言模型(LLMs)通过直接提示或微调实现了更复杂的审核能力,但其泛化性、可解释性和对未知/模糊案例的适应性仍不足。本文提出一种新型审核框架CHAIRO,利用类比示例增强规则归纳与决策可靠性。该方法端到端优化类比检索、规则生成与审核分类,实现审核规则对多样内容场景的动态适配。大量实验表明,该方法在审核准确率和规则质量上均显著优于注入规则的微调基线及多阶段静态RAG管道。进一步的人工评估和外部模型泛化测试证实,该框架生成的规则具有更高清晰度、可解释性和适用性。结果表明,基于类比示例的方法可推动真实应用中更鲁棒、可解释、可泛化的內容审核。

原文摘要 · Abstract (English)

Content moderation in online platforms faces persistent challenges due to the evolving complexity of user-generated content and the limitations of traditional rule-based and machine learning approaches. While recent advances in large language models (LLMs) have enabled more sophisticated moderation via direct prompting or fine-tuning, these approaches often exhibit limited generalization, interpretability, and adaptability to unseen or ambiguous cases. In this work, we propose a novel moderation framework that leverages analogical examples to enhance rule induction and decision reliability. Our approach integrates end-to-end optimization of analogical retrieval, rule generation, and moderation classification, enabling the dynamic adaptation of moderation rules to diverse content scenarios. Through comprehensive experiments, we demonstrate that our method significantly outperforms both rule-injected fine-tuning baselines and multi-stage static RAG pipelines in terms of moderation accuracy and rule quality. Further evaluations, including human assessments and external model generalization tests, confirm that our framework produces rules with better clarity, interpretability, and applicability. These findings show that analogical example-driven methods can advance robust, explainable, and generalizable content moderation in real-world applications.

内容审核类比推理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。