提升代码补全评估与数据构建,更贴合开发者真实体验。
Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion
- 从概率建模出发设计新评估指标LCP和ROUGE-LCP。
- 在用户感知一致性上优于传统指标,模型性能显著提升。
- 构建结构保持的代码图,增强跨模块依赖理解。
基于大语言模型的代码补全技术显著提升了程序员的开发效率。然而,在实际应用中,当前常用的代码补全评估指标与用户真实感知之间仍存在差距。为解决这一问题,本文从概率建模角度提出两种新的代码补全评估指标——LCP和ROUGE-LCP。此外,针对大语言模型在仓库级代码补全场景中缺乏有效的结构语义建模和跨模块依赖信息的问题,提出一种基于结构保持与语义重排代码图(SPSR-Graph)的数据处理方法。通过理论分析与实验验证,证明所提评估指标在用户感知一致性方面具有优势,且该数据处理方法能有效提升模型性能。
原文摘要 · Abstract (English)
Code completion technology based on large language model has significantly improved the development efficiency of programmers. However, in practical applications, there remains a gap between current commonly used code completion evaluation metrics and users' actual perception. To address this issue, we propose two evaluation metrics for code completion tasks--LCP and ROUGE-LCP, from the perspective of probabilistic modeling. Furthermore, to tackle the lack of effective structural semantic modeling and cross-module dependency information in LLMs for repository-level code completion scenarios, we propose a data processing method based on a Structure-Preserving and Semantically-Reordered Code Graph (SPSR-Graph). Through theoretical analysis and experimental validation, we demonstrate the superiority of the proposed evaluation metrics in terms of user perception consistency, as well as the effectiveness of the data processing method in enhancing model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。