为大模型对韩文化对齐设计了正向引导机制,提升文化合规性。
Korean Culture into LLM Alignment: Toward Cultural Coherence
- 基于韩语法律与社会规范构建文化安全响应指南
- 在六款开源大模型上提升韩文化安全率,通用能力无明显下降
- 可生成引用韩国法律条文的建设性回应,适合本地化应用
当前大模型的文化对齐研究多聚焦于抑制不当输出,我们主张还需构建积极的文化一致性定义。本文针对韩国文化,设计了一套以提示驱动的模型生成器为核心的对齐数据流水线,扩展了韩国有害内容分类体系,并以韩式法律框架、社会规范和解释惯例为核心制定安全响应政策。三款前沿模型在该政策指导下生成候选回复,经对比偏好微调(DPO)后,在六款开源大模型中显著提升了韩文化安全率,同时在韩语通用能力基准上未出现明显退化。定性分析显示,微调后模型能准确引用韩国法律条文与制度流程,并在适当情境下提供建设性文化背景信息。
原文摘要 · Abstract (English)
Cultural-aspect work on large language models is dominated by a negative target: which outputs to suppress. We argue that a constructive counterpart is also needed, a working definition of what a culturally coherent response is rather than only what it must avoid, and instantiate it for Korean. We design an alignment-data pipeline around a prompt-based LLM seed generator that expands a Korean harm taxonomy, with a Korean-culturally-adapted safe-response policy at its centre: a per-category guideline grounded in Korean legal frameworks, social norms, and interpretive conventions, against which three frontier models each produce a candidate response. DPO fine-tuning on the resulting triplets improves the Korean cultural safe rate across six open-weight LLMs while causing no large degradation on Korean general-capability benchmarks, and qualitative outputs show fine-tuned models naming Korean statutes and institutional procedures and, where appropriate, supplying constructive Korean-context information alongside refusal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。