综述大模型在多语言代码智能中的应用与挑战
Large Language Models for Multilingual Code Intelligence: A Survey

- 聚焦多语言代码生成与语义保持的代码翻译任务
- 指出当前模型对低资源语言如Rust表现较弱
- 适合关注跨语言编程与AI辅助开发的研究者
大语言模型已推动AI辅助软件工程变革,但现有研究仍偏向高资源语言(如Python),在Rust、OCaml等语言上性能较弱。由于真实系统普遍存在多语言混合特性,构建可靠的多语言代码智能至关重要。本综述聚焦两大核心任务:基于统一自然语言需求的多语言代码生成,以及保持语义一致性的多语言代码翻译。文章梳理代表性方法、基准测试与评估指标,揭示可信跨语言泛化的挑战与机遇。
原文摘要 · Abstract (English)
Large language models have transformed AI-assisted software engineering, but current research remains biased toward high-resource languages such as Python, with weaker performance in languages like Rust and OCaml. Since real-world systems are inherently polyglot, robust multilingual code intelligence is crucial. This survey focuses on two key tasks: multilingual code generation from shared natural-language requirements, and multilingual code translation that preserves semantics across languages. It reviews representative methods, benchmarks, and evaluation metrics, and highlights challenges and opportunities for trustworthy cross-language generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。