arXiv:2512.12501cs.AI2025-12

让文生图更安全,自动过滤有害提示并生成合规图像

SafeGen: Embedding Ethical Safeguards in Text-to-Image Generation

  • 用微调文本分类器和优化扩散模型双模块实现伦理防护
  • 生成图像质量高(IS=3.52),识别有害提示准确率81%
  • 适合教育、内容创作等需兼顾创意与合规的场景

生成式人工智能为创意表达、教育和研究带来新机遇,但文生图系统如DALL·E、Stable Diffusion和Midjourney也引发双重使用风险:放大社会偏见、生成高保真虚假信息、侵犯知识产权。本文提出SafeGen框架,将伦理防护嵌入生成流程,基于可信AI原则设计。该框架包含两个互补组件:基于英越双语数据集与公平性训练的BGE-M3文本分类器,用于过滤有害或误导性提示;以及优化后的Hyper-SD扩散模型,生成高保真且语义对齐的图像。定量评估显示,Hyper-SD在图像质量(IS=3.52)、分布相似性(FID=22.08)和结构相似性(SSIM=0.79)上表现优异,BGE-M3的F1分数达0.81。消融实验验证了领域微调对两模块的关键作用。案例研究证明其在阻止危险提示、生成包容性教学材料和维护学术诚信方面的实际价值。

原文摘要 · Abstract (English)

Generative Artificial Intelligence (AI) has created unprecedented opportunities for creative expression, education, and research. Text-to-image systems such as DALL.E, Stable Diffusion, and Midjourney can now convert ideas into visuals within seconds, but they also present a dual-use dilemma, raising critical ethical concerns: amplifying societal biases, producing high-fidelity disinformation, and violating intellectual property. This paper introduces SafeGen, a framework that embeds ethical safeguards directly into the text-to-image generation pipeline, grounding its design in established principles for Trustworthy AI. SafeGen integrates two complementary components: BGE-M3, a fine-tuned text classifier that filters harmful or misleading prompts, and Hyper-SD, an optimized diffusion model that produces high fidelity, semantically aligned images. Built on a curated multilingual (English- Vietnamese) dataset and a fairness-aware training process, SafeGen demonstrates that creative freedom and ethical responsibility can be reconciled within a single workflow. Quantitative evaluations confirm its effectiveness, with Hyper-SD achieving IS = 3.52, FID = 22.08, and SSIM = 0.79, while BGE-M3 reaches an F1-Score of 0.81. An ablation study further validates the importance of domain-specific fine-tuning for both modules. Case studies illustrate SafeGen's practical impact in blocking unsafe prompts, generating inclusive teaching materials, and reinforcing academic integrity.

文生图伦理安全扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。