SMILE融合词义与关键词匹配,更准确评估问答系统表现。
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- 结合句子级语义与关键词级语义,兼顾语义与精确匹配。
- 在文本、图像和视频问答任务中与人工评分高度相关。
- 轻量高效,避免大模型的高成本与不一致问题。
传统问答评估指标如ROUGE、METEOR和精确匹配(EM)过度依赖n元语法级别的词汇相似性,常忽略深层语义理解。尽管BERTScore和MoverScore利用上下文嵌入改善语义评估,但仍缺乏对句子级与关键词级语义的灵活权衡,且忽视了仍重要的词汇相似性。基于大语言模型的评估器虽强大,但存在成本高、偏见、不一致及幻觉等问题。为此,我们提出SMILE:一种融合词汇精确性的语义评估新方法,结合句子级语义理解、关键词级语义理解与关键词直接匹配,实现词汇精确性与语义相关性的平衡。在文本、图像和视频问答任务上的广泛基准测试表明,SMILE与人工评分高度相关,且计算开销低,有效弥合了词汇与语义评估之间的鸿沟。
原文摘要 · Abstract (English)
Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。