检验解题器耗时是否反映人类感知难度,发现二者不一致。
Evaluating SAT Solver Metrics as Predictors of Human-Perceived Nonogram Difficulty

- 将数独类谜题建模为约束满足问题,用SAT求解器评估复杂度。
- 用户研究显示解题器指标与人感难度无显著相关性。
- 专家与新手解题策略不同,人类偏好复杂推理而非简单搜索。
算法求解耗时常被默认等同于人类感知的谜题难度,但这一假设极少通过真实用户数据验证。本文针对流行的逻辑谜题非诺格拉姆(Nonograms)进行检验,将其建模为约束满足问题并使用现有SAT求解器求解。通过用户实验,收集参与者的行为数据与主观难度评分。结果表明,无论是主观难度还是行为信号,均与SAT求解器的指标无显著相关性;但发现个体经验水平调节了解题器复杂度与主观难度之间的关系。在此过程中,识别出若干反复出现的人类解题策略,表明人类更偏好复杂的传播推理,与求解器衡量的复杂度存在差异。
原文摘要 · Abstract (English)
Algorithmic solver effort is often assumed to align with perceived puzzle difficulty, but this assumption is rarely tested against human solving data. We evaluate this assumption for Nonograms, a popular logic puzzle similar to Sudoku in which numeric clues along each row and column determine a unique solution grid. We formulate Nonograms as a constraint satisfaction problem and solve them using existing SAT solvers. We then conduct a user study in which we collect data on both participant interactions and reported difficulty. We find that neither participants' reported difficulty nor their behavioural signals correlate meaningfully with SAT solver metrics; however, we find evidence that expertise moderates the relationship between solver metrics and reported difficulty. In this process, we uncover distinct, recurring solving strategies that indicate human preference for complex propagation, diverging from solver-measured complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。