针对车载助手的韩语礼貌用语精准控制难题,提出细粒度评估框架。
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control

- 构建面向韩语礼貌等级的细粒度评估体系
- 发现现有模型在礼貌用语上表现不稳定,对话主动性弱
- 适合关注车载语音交互安全与本地化适配的研究者
随着大语言模型在车载对话系统中的广泛应用,因缺乏针对实际部署需求的领域专用评估标准,选择最优模型仍具挑战。本文提出一种面向车载助手的新评估框架,重点关注韩语本地化问题。实证分析揭示显著规律:当前大模型在韩语礼貌等级的精细控制上仍不稳定,表明本地化场景中需对话语级表达进行显式评估;同时,模型在澄清与主动引导等策略性对话指标上表现较弱,这源于任务本身的主观复杂性。因此,本框架采用保守评估立场以优先保障可靠性。研究强调,汽车AI应从通用能力转向精准语言定制与可靠、安全导向的交互管理。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) are increasingly integrated into in-vehicle conversational systems, identifying the optimal model remains challenging due to the lack of domain-specific evaluation standards tailored to real-world deployment requirements. In this paper, we propose a novel evaluation framework for in-vehicle assistants, with a particular focus on Korean-language localization. Our empirical analysis reveals notable patterns in model behavior. First, fine-grained Korean honorific control remains unstable in current LLMs, indicating that precise speech-level realization must be explicitly evaluated in localization settings. Second, models exhibit weaker performance in strategic conversational metrics like clarification and proactivity. Our analysis suggests this stems from the inherent subjective complexity of these tasks, where our framework adopts a conservative evaluation stance to prioritize reliability. Together, our findings underscore that automotive AI must move beyond general competence toward precise linguistic tailoring and reliable, safety-oriented interaction management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。