arXiv:2602.13860cs.CL2026-02中稿 · the PhD Symposium …

让大模型更懂领域、更安全、更符合文化差异。

Tutoring Large Language Models to be Domain-adaptive, Precise, and Safe

  • 用人类反馈和偏好建模提升模型的社会语言能力。
  • 在推理阶段对齐安全策略,减少有害输出。
  • 适合需要高精度与跨文化适配的落地场景。

本研究致力于构建一种「负责任智能」框架,以调和大型语言模型(LLMs)强大的生成能力与实际部署中的严格要求之间的矛盾。随着大模型成为人工智能变革的核心,亟需从通用架构转向具备上下文感知、内在安全性和全球文化敏感性的系统。研究聚焦三个相互关联的维度:通过领域自适应确保技术精准性,通过伦理严谨性降低对抗性漏洞,通过文化与多语言对齐促进全球包容性。方法路径从传统的监督微调转向推理时对齐以保障安全,最终借助人类反馈与偏好建模实现社会语言敏锐度。

原文摘要 · Abstract (English)

The overarching research direction of this work is the development of a ''Responsible Intelligence'' framework designed to reconcile the immense generative power of Large Language Models (LLMs) with the stringent requirements of real-world deployment. As these models become a transformative force in artificial intelligence, there is an urgent need to move beyond general-purpose architectures toward systems that are contextually aware, inherently safer, and deeply respectful of global cultural nuances. This research navigates three interconnected threads: domain adaptation to ensure technical precision, ethical rigor to mitigate adversarial vulnerabilities, and cultural/multilingual alignment to promote global inclusivity. The methodological trajectory moves from classical supervised adaptation for task-specific demands to decoding-time alignment for safety, finally leveraging human feedback and preference modeling to achieve sociolinguistic acuity.

大模型安全文化对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。