arXiv:2606.19344cs.CLcs.AI2026-06

通过可视化生成路径聚合,揭示大模型中隐藏的偏见。

Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation

论文配图:Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation
图 1 · 摘自论文原文
  • 构建树状结构聚合数百次随机生成结果,映射潜在语义分支。
  • 对比不同上下文下反事实词概率,发现代词压制等隐蔽偏见。
  • 适合研究者和伦理审查员分析模型系统性偏差,降低认知负担。

大型语言模型(LLMs)存在难以评估的表征与句法偏见,因其文本生成具有随机性。传统审计方法依赖单次输出或静态指标,无法捕捉低概率生成路径中的隐含偏见。本文提出TreeTracer,一种基于系统性扰动分析的视觉分析工具:通过替换输入提示中的本体定义术语,将数百次随机生成结果聚合为语法对齐的层次化结构,并利用辅助语言模型进行分类感知的节点合并,最终以自定义桑基图可视化。通过并列两个本体驱动的树结构,可直接比较不同语义上下文,支持系统性偏见检测。由于可视化仅反映模型部分行为,系统进一步采用对比推理计算并直观展示跨上下文的反事实词概率,降低误判风险。案例研究对比了未对齐的GPT-2 XL与宪法对齐的Apertus模型,成功揭示了代词抑制、对话边缘化等隐藏表征伤害。初步用户研究证实,聚合对比界面显著降低认知负荷,有效辅助分析师识别系统性偏见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generation. Standard auditing methods rely on a single output inspection or static automated metrics. These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability generation branches. This paper introduces TreeTracer, a visual analytics tool designed to evaluate LLM bias through aggregated comparison. Using a systematic perturbation analysis pipeline, the tool replaces ontology-defined terms in each input prompt, aggregates hundreds of stochastic generations into a syntax-aligned hierarchical structure, and then performs classification-aware node merging with an auxiliary language model. The resulting structure is visualized through a custom Sankey diagram. By juxtaposing two ontology-driven trees, the workspace enables direct comparison between semantic contexts and supports systematic bias detection. Because any visualization reflects only a subset of the model's learned behavior, the system further applies contrastive inference to compute and directly display counterfactual token probabilities across contexts, reducing the risk of misinterpreting the presence of bias. We validate the workspace through case studies comparing an unaligned baseline model GPT-2 XL against the constitutionally aligned Apertus models. The visual aggregation successfully exposes hidden representational harms, such as counterfactual pronoun suppression and conversational marginalization of individuals. A preliminary user study confirms that the aggregated comparative interface reduces cognitive load and effectively supports analysts in detecting systemic biases.

模型偏见可视化分析生成路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。