arXiv:2410.20811cs.LGcs.AI2024-10NAACL被引 9

用概念引导生成更可信的国际象棋解说,兼顾准确与流畅。

Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation

  • 结合专家模型决策力与大模型语言能力,按概念优先级生成解说。
  • 生成内容经人工与自动评估,准确率、信息量和流畅度均显著提升。
  • 适合需要可解释性的人机协作场景,如教育或对弈分析。

基于深度学习的专家模型在国际象棋等决策领域已达到超人水平,但对其决策过程的解释仍缺乏研究,而解释对模型可解释性和人类教育至关重要。专家模型输出精准却难懂,大语言模型虽能生成流畅解说,却因决策能力有限易产生幻觉。为弥合这一差距,本文以国际象棋解说为例,同时解决生成与评估问题。提出概念引导的国际象棋解说生成(CCC),融合专家模型的决策优势与大模型的语言优势,通过优先级化的概念解释实现高质量输出;提出基于GPT的国际象棋解说评估(GCC-Eval),利用专家知识从信息量与语言质量双维度评估解说。实验结果经人工评审与自动评估双重验证,表明CCC生成的解说兼具准确性、信息丰富性与语言流畅性。

原文摘要 · Abstract (English)

Deep learning-based expert models have reached superhuman performance in decision-making domains such as chess and Go. However, it is under-explored to explain or comment on given decisions although it is important for model explainability and human education. The outputs of expert models are accurate, but yet difficult to interpret for humans. On the other hand, large language models (LLMs) can produce fluent commentary but are prone to hallucinations due to their limited decision-making capabilities. To bridge this gap between expert models and LLMs, we focus on chess commentary as a representative task of explaining complex decision-making processes through language and address both the generation and evaluation of commentary. We introduce Concept-guided Chess Commentary generation (CCC) for producing commentary and GPT-based Chess Commentary Evaluation (GCC-Eval) for assessing it. CCC integrates the decision-making strengths of expert models with the linguistic fluency of LLMs through prioritized, concept-based explanations. GCC-Eval leverages expert knowledge to evaluate chess commentary based on informativeness and linguistic quality. Experimental results, validated by both human judges and GCC-Eval, demonstrate that CCC generates commentary which is accurate, informative, and fluent.

可解释AI自然语言生成国际象棋大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。