PromptGuard通过智能提示框架,为弱势群体生成更安全、公平的文本。
PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
- 设计混合提示技术VulnGuard,结合真实数据与推理链防止有害输出。
- 理论证明可降低25%-30%的伤害风险,基于熵边界与帕累托最优。
- 适合关注AI伦理、安全生成的开发者与政策制定者使用。
大型语言模型在实际应用中对包括LGBTQ+群体、单亲家庭及边缘化社区在内的弱势人群生成有害、偏见或误导性信息的风险日益加剧。现有安全方法依赖事后过滤或通用对齐技术,无法从源头主动预防。本文提出PromptGuard——一种模块化提示框架,核心创新为VulnGuard提示:一种基于真实世界数据驱动的对比学习混合技术,可有效阻止有害内容生成。该技术融合精选GitHub仓库中的少样本示例、伦理思维链推理与自适应角色提示,构建针对特定人群的保护屏障。框架采用理论多目标优化,通过形式化证明,在熵边界与帕累托最优下实现25%-30%的分析性伤害降低。PromptGuard整合六项核心模块:输入分类、VulnGuard提示、伦理原则集成、外部工具交互、输出验证与用户-系统交互,形成实时危害预防的智能专家系统。研究提供完整的数学形式化,包含收敛性证明、基于信息论的漏洞分析以及使用GitHub来源数据集的理论验证框架,为系统性实证研究奠定数学基础。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) in real-world applications poses unprecedented risks of generating harmful, biased, or misleading information to vulnerable populations including LGBTQ+ individuals, single parents, and marginalized communities. While existing safety approaches rely on post-hoc filtering or generic alignment techniques, they fail to proactively prevent harmful outputs at the generation source. This paper introduces PromptGuard, a novel modular prompting framework with our breakthrough contribution: VulnGuard Prompt, a hybrid technique that prevents harmful information generation using real-world data-driven contrastive learning. VulnGuard integrates few-shot examples from curated GitHub repositories, ethical chain-of-thought reasoning, and adaptive role-prompting to create population-specific protective barriers. Our framework employs theoretical multi-objective optimization with formal proofs demonstrating 25-30% analytical harm reduction through entropy bounds and Pareto optimality. PromptGuard orchestrates six core modules: Input Classification, VulnGuard Prompting, Ethical Principles Integration, External Tool Interaction, Output Validation, and User-System Interaction, creating an intelligent expert system for real-time harm prevention. We provide comprehensive mathematical formalization including convergence proofs, vulnerability analysis using information theory, and theoretical validation framework using GitHub-sourced datasets, establishing mathematical foundations for systematic empirical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。