融合多个评估指标,提升大模型输出可信度的跨领域评价能力
Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
- 用树模型融合基础可信度指标,根据人工判断确定权重
- 新指标与人工评判相关性在各领域均显著提升
- 提供统一数据集和评测框架,支持跨领域复现
本文提出一种改进大语言模型(LLM)可信度评估准确性的方法。该方法通过将基础可信度指标组合成融合指标,提升对模型输出可信度的评估效果。融合策略采用树模型识别各指标重要性,其训练基于人工对模型响应可信度的评判。实验表明,该融合指标在所有测试领域中与人工评判的相关性均更强。提升可信度评估能力,有助于增强对模型的信任,推动其在更广泛场景中的应用。此外,本文还整合了问答与对话领域的多个数据集,构建统一评估环境,并纳入人工评判与模型输出,支持跨领域可信度评估的复现与测试。
原文摘要 · Abstract (English)
We present a methodology for improving the accuracy of faithfulness evaluation in Large Language Models (LLMs). The proposed methodology is based on the combination of elementary faithfulness metrics into a combined (fused) metric, for the purpose of improving the faithfulness of LLM outputs. The proposed strategy for metric fusion deploys a tree-based model to identify the importance of each metric, which is driven by the integration of human judgements evaluating the faithfulness of LLM responses. This fused metric is demonstrated to correlate more strongly with human judgements across all tested domains for faithfulness. Improving the ability to evaluate the faithfulness of LLMs, allows for greater confidence to be placed within models, allowing for their implementation in a greater diversity of scenarios. Additionally, we homogenise a collection of datasets across question answering and dialogue-based domains and implement human judgements and LLM responses within this dataset, allowing for the reproduction and trialling of faithfulness evaluation across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。