arXiv:2606.16368cs.CLcs.LG2026-06

用语义约束验证框架,精准评估大模型个性化表现。

Evaluating LLM Personalization via Semantic Constraint Verification

论文配图:Evaluating LLM Personalization via Semantic Constraint Verification
图 1 · 摘自论文原文
  • 将语义映射为真值集,通过NLI模型验证个性化约束
  • 识别出四种行为模式,准确率接近人工标注
  • 可定位关键句子,结果可解释且效率提升2100倍

当前大语言模型个性化评估依赖脆弱的表面匹配指标或计算成本高昂的LLM作为裁判,均缺乏可解释性。为此,我们提出自然语言推理约束验证(NLICV),一种可扩展、语义不变的框架,将句子意义映射为真值集,通过自然语言推理(NLI)模型验证个性化约束。超越二元评分,NLICV将模型行为分为四类:个性化、泛化、奉承和失败。大量实验表明,NLICV与人工标注高度一致,同时大幅降低LLM裁判带来的延迟和令牌消耗(最高达2100倍加速)。最后,通过消融分析,NLICV能精确定位驱动约束验证的关键句子,提供真实可信、可理解的评估证据。

原文摘要 · Abstract (English)

Current evaluation paradigms for Large Language Model (LLM) personalization rely heavily on brittle surface-matching metrics or computationally expensive LLM-as-a-judge protocols, both of which lack interpretability. To address these limitations, we introduce Natural Language Inference Constraint Verification (NLICV), a scalable, semantically invariant framework that maps sentence meanings to truth-condition sets to verify personalization constraints via a Natural Language Inference (NLI) model. Moving beyond binary scoring, NLICV categorizes LLM behaviors into four distinct modes: personalization, generalization, sycophancy, and failure. Extensive experiments demonstrate that NLICV aligns closely with human annotations while drastically reducing the latency and token costs associated with LLM judges (up to 2100 inference speedup). Finally, through an ablation-based procedure, NLICV pinpoints the exact sentences driving the constraint verification, yielding faithful, understandable evidence for its evaluations.

个性化评估语义验证可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。