arXiv:2410.02381cs.CLcs.AI2024-10ICLR被引 20

用人类偏好校准生成任务评价指标,让评分更贴近真实判断。

MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

  • 基于人类偏好监督优化多个指标组合,提升整体评估能力。
  • 在多语言、多领域场景下显著提升与人类判断的一致性。
  • 可灵活扩展,轻松集成到各类生成任务评估中。

理解性能评估指标的质量对于确保模型输出符合人类偏好至关重要。然而,现有指标往往在某些维度表现优异,却难以覆盖人类偏好的多样性。为此,必须系统性地将指标校准至人类偏好的特定方面。我们提出 MetaMetrics,一种通过监督方式校准的元指标,用于跨模态生成任务的评估。该方法通过优化现有指标的组合,增强其与人类偏好的对齐程度。实验表明,MetaMetrics 在语言与视觉下游任务中均表现出灵活性与有效性,在多语言和多领域场景下均有显著提升。其与人类偏好高度一致,具备强可扩展性,且易于集成到任意应用中。这使其成为提升生成任务评估质量的强大工具,确保指标在多样环境中更真实反映人类判断。

原文摘要 · Abstract (English)

Understanding the quality of a performance evaluation metric is crucial for ensuring that model outputs align with human preferences. However, it remains unclear how well each metric captures the diverse aspects of these preferences, as metrics often excel in one particular area but not across all dimensions. To address this, it is essential to systematically calibrate metrics to specific aspects of human preference, catering to the unique characteristics of each aspect. We introduce MetaMetrics, a calibrated meta-metric designed to evaluate generation tasks across different modalities in a supervised manner. MetaMetrics optimizes the combination of existing metrics to enhance their alignment with human preferences. Our metric demonstrates flexibility and effectiveness in both language and vision downstream tasks, showing significant benefits across various multilingual and multi-domain scenarios. MetaMetrics aligns closely with human preferences and is highly extendable and easily integrable into any application. This makes MetaMetrics a powerful tool for improving the evaluation of generation tasks, ensuring that metrics are more representative of human judgment across diverse contexts.

生成评估人类偏好指标校准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。