通过用户偏好对比,自动选出逻辑谜题中更易懂的解题步骤。
Preference Elicitation for Step-Wise Explanations in Logic Puzzles
- 用成对比较学习用户偏好,动态归一化多尺度评价指标。
- 引入新型查询策略MACHOP,提升解释生成多样性与质量。
- 在数独和逻辑网格谜题上验证,比传统方法更优,适合人机交互场景。
分步解释可通过逐步展示约束条件推导决策变量取值,来说明逻辑谜题等求解问题。然而,存在大量候选解释步骤,其包含的约束集与推导决策各不相同。为识别最易理解的步骤,需定义用户导向的目标函数以量化每一步的质量。但设计有效目标函数具有挑战性。本文借鉴机器学习中的交互式偏好获取方法,利用成对比较学习用户偏好。针对解释质量涉及多个量级差异显著的子目标,提出两种动态归一化技术以稳定学习过程。同时发现生成的比较中存在大量相似解释,为此提出MACHOP(Multi-Armed CHOice Perceptron)——一种融合非支配约束与置信上限驱动多样化的查询生成策略。在人工用户与真实用户评估中,分别对数独与逻辑网格谜题进行测试,结果表明MACHOP始终优于标准方法,显著提升解释质量。
原文摘要 · Abstract (English)
Step-wise explanations can explain logic puzzles and other satisfaction problems by showing how to derive decisions step by step. Each step consists of a set of constraints that derive an assignment to one or more decision variables. However, many candidate explanation steps exist, with different sets of constraints and different decisions they derive. To identify the most comprehensible one, a user-defined objective function is required to quantify the quality of each step. However, defining a good objective function is challenging. Here, interactive preference elicitation methods from the wider machine learning community can offer a way to learn user preferences from pairwise comparisons. We investigate the feasibility of this approach for step-wise explanations and address several limitations that distinguish it from elicitation for standard combinatorial problems. First, because the explanation quality is measured using multiple sub-objectives that can vary a lot in scale, we propose two dynamic normalization techniques to rescale these features and stabilize the learning process. We also observed that many generated comparisons involve similar explanations. For this reason, we introduce MACHOP (Multi-Armed CHOice Perceptron), a novel query generation strategy that integrates non-domination constraints with upper confidence bound-based diversification. We evaluate the elicitation techniques on Sudokus and Logic-Grid puzzles using artificial users, and validate them with a real-user evaluation. In both settings, MACHOP consistently produces higher-quality explanations than the standard approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。