区分大模型输出的结构与意图真实度,发现多数完美表现背后存在意图缺失。
Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation
- 按语义维度拆解评估模型是否忠实还原用户意图。
- 中文输出25.7%、英文输出58.6%在整体评分高但存在意图偏差。
- 该方法更贴近人类判断,适合精细化评估用户定制任务。
整体评估分数反映输出质量总体水平,但无法区分模型是否准确复现了用户的结构要求或具体意图。本文提出一种基于结构化提示消融实验的维度级意图保真度评估框架,覆盖2,880条跨三语言、三任务领域的输出,对每个语义维度分别测量结构恢复与意图保真度。结果揭示系统性结构-意图分离现象:在具有完整配对评分的中文输出中,25.7%获得满分(GA=5)但存在可测量的维度意图缺陷;英文输出中该比例达58.6%。人工评估证实这些“分裂区”输出确实存在质量问题,且维度保真度评分比整体评分更可靠地匹配人类判断。对2,520个消融单元进行公私成分分解,刻画了模型在缺失意图时的补偿能力;代理标注区分了先验可推断性与默认可恢复性。权重扰动实验表明,适度失配通常被吸收,而严重维度反转则始终有害。研究证明,维度级意图保真度评估是用户特定任务中不可或缺的补充手段。
原文摘要 · Abstract (English)
Holistic evaluation scores capture overall output quality but do not distinguish whether a model reproduced the structural form of a user's request from whether it preserved the user's specific intent. We propose a dimension-level intent fidelity evaluation framework, applied here through a structured prompt ablation study across 2,880 outputs spanning three languages, three task domains, and six LLMs, that separately measures structural recovery and intent fidelity for each semantic dimension. This framework reveals a systematic structural-fidelity split: among Chinese-language outputs with complete paired scores, 25.7% received perfect holistic alignment scores (GA=5) while exhibiting measurable dimensional intent deficits; among English-language outputs, this proportion rose to 58.6%. Human evaluation confirmed that these split-zone outputs represent genuine quality deficits and that dimensional fidelity scores track human judgements more reliably than holistic scores do. A public-private decomposition of 2,520 ablation cells characterises when models successfully compensate for missing intent and when they fail, while proxy annotation distinguishes prior inferability from default recoverability. A weight-perturbation experiment shows that moderate misalignment is typically absorbed, whereas severe dimensional inversion is consistently harmful. These findings demonstrate that dimension-level intent fidelity evaluation is a necessary complement to holistic assessment when evaluating LLM outputs for user-specific tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。