让大模型更懂领域、更安全、更符合文化差异。
Tutoring Large Language Models to be Domain-adaptive, Precise, and Safe
- 用人类反馈和偏好建模提升模型的社会语言能力。
- 在推理阶段对齐安全策略,减少有害输出。
- 适合需要高精度与跨文化适配的落地场景。
本研究致力于构建一种「负责任智能」框架,以调和大型语言模型(LLMs)强大的生成能力与实际部署中的严格要求之间的矛盾。随着大模型成为人工智能变革的核心,亟需从通用架构转向具备上下文感知、内在安全性和全球文化敏感性的系统。研究聚焦三个相互关联的维度:通过领域自适应确保技术精准性,通过伦理严谨性降低对抗性漏洞,通过文化与多语言对齐促进全球包容性。方法路径从传统的监督微调转向推理时对齐以保障安全,最终借助人类反馈与偏好建模实现社会语言敏锐度。
原文摘要 · Abstract (English)
The overarching research direction of this work is the development of a ''Responsible Intelligence'' framework designed to reconcile the immense generative power of Large Language Models (LLMs) with the stringent requirements of real-world deployment. As these models become a transformative force in artificial intelligence, there is an urgent need to move beyond general-purpose architectures toward systems that are contextually aware, inherently safer, and deeply respectful of global cultural nuances. This research navigates three interconnected threads: domain adaptation to ensure technical precision, ethical rigor to mitigate adversarial vulnerabilities, and cultural/multilingual alignment to promote global inclusivity. The methodological trajectory moves from classical supervised adaptation for task-specific demands to decoding-time alignment for safety, finally leveraging human feedback and preference modeling to achieve sociolinguistic acuity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。