arXiv:2505.14585cs.CL2025-05EMNLP被引 7

用强化学习让大模型在隐私安全场景下更懂上下文推理。

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

  • 基于上下文完整性理论,用规则奖励引导模型做合规推理。
  • 安全隐私基准准确率提升8.58%,通用推理能力也显著增强。
  • 适合关注法律合规与模型安全的AI研发者使用。

大型语言模型虽能力强大,但伴随显著的安全与隐私风险。现有缓解策略常依赖敏感模式匹配,忽视上下文推理能力,且未遵循权威标准,导致系统性合规风险。为此,本文基于上下文完整性(CI)理论,将安全与隐私问题建模为情境化合规问题,并对齐GDPR、EU AI Act和HIPAA三大监管标准。采用基于规则奖励的强化学习方法,激励模型在高风险场景中保持上下文推理能力。实验表明,该方法不仅使安全/隐私基准准确率提升8.58%,还显著增强通用推理能力:在OpenThinker-7B模型上,MMLU和LegalBench分别取得+2.05%和+8.98%的准确率提升,优于其基线模型Qwen2.5-7B-Instruct。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) exhibit remarkable capabilities, they also introduce significant safety and privacy risks. Current mitigation strategies often fail to preserve contextual reasoning capabilities in risky scenarios. Instead, they rely heavily on sensitive pattern matching to protect LLMs, which limits the scope. Furthermore, they overlook established safety and privacy standards, leading to systemic risks for legal compliance. To address these gaps, we formulate safety and privacy issues into contextualized compliance problems following the Contextual Integrity (CI) theory. Under the CI framework, we align our model with three critical regulatory standards: GDPR, EU AI Act, and HIPAA. Specifically, we employ reinforcement learning (RL) with a rule-based reward to incentivize contextual reasoning capabilities while enhancing compliance with safety and privacy norms. Through extensive experiments, we demonstrate that our method not only significantly enhances legal compliance (achieving a +8.58% accuracy improvement in safety/privacy benchmarks) but also further improves general reasoning capability. For OpenThinker-7B, a strong reasoning model that significantly outperforms its base model Qwen2.5-7B-Instruct across diverse subjects, our method enhances its general reasoning capabilities, with +2.05% and +8.98% accuracy improvement on the MMLU and LegalBench benchmark, respectively.

大模型安全强化学习合规推理隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。