让语言模型的隐性决策可见可调,帮用户看清问题背后的思维路径。
Navigating the Conceptual Multiverse

- 构建概念多宇宙系统,可视化并可干预模型对问题的框架与价值判断。
- 在哲学、对齐标注、诗歌创作三领域中提升用户对问题本质的理解深度。
- 提供可验证的推理规范,确保决策结构清晰、完整且符合专业认知。
当语言模型回答开放性问题时,会做出隐性的概念决策,影响输出结果,但用户往往无法理解这些决策过程。本文借鉴统计学中的多宇宙分析思想,提出并评估了一个名为‘概念多宇宙’的交互式系统,用于呈现如问题如何定义、何为重要等概念性选择,使用户能够透明查看、主动修改,并依据领域内合理推理进行检验。为确保该结构不误导而具有实用性,我们设计了一套通用验证框架,通过专家级推理校准良好决策结构应具备的明确性与完整性。在三个不同领域(哲学、对齐标注、诗歌创作)的应用中,参与者借助该系统建立了更深入的问题认知地图:哲学学生重写论文时确立了更清晰的论述框架并反转了原有论点;对齐标注者从表面偏好转向基于用户意图与伤害风险的深层推理;诗人识别出创作中的模式,进而明晰自身审美偏好。
原文摘要 · Abstract (English)
When language models answer open-ended problems, they implicitly make hidden decisions that shape their outputs, leaving users with uncontextualized answers rather than a working map of the problem; drawing on multiverse analysis from statistics, we build and evaluate the conceptual multiverse, an interactive system that represents conceptual decisions such as how to frame a question or what to value as a space users can transparently inspect, intervenably change, and check against principled domain reasoning; for this structure to be worth navigating rather than misleading, it must be rigorous and checkable against domain reasoning norms, so we develop a general verification framework that enforces properties of good decision structures like unambiguity and completeness calibrated by expert-level reasoning; across three domains, the conceptual multiverse helped participants develop a working map of the problem, with philosophy students rewriting essays with sharper framings and reversed theses, alignment annotators moving from surface preferences to reasoning about user intent and harm, and poets identifying compositional patterns that clarified their taste.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。