arXiv:2601.04207cs.CLcs.AI2026-01

用轻量方法纠正大模型在社交媒体中的意识形态偏差。

Ideology as a Problem: Lightweight Logit Steering for Annotator-Specific Alignment in Social Media Analysis

  • 通过内部特征计算偏见分数,直接调整输出概率。
  • 无需微调模型,即可对齐特定用户观点,效果可量化。
  • 适合需要快速适配用户立场的社交内容分析场景。

大语言模型在内部以低维结构组织政治意识形态,但与人类意识形态空间存在系统性偏差,且因模型而异、可测量。本文提出一种轻量级线性探测器,既能量化该偏差,又能最小化修正输出层。通过分析模型内部特征计算偏见分数,并直接调整最终输出概率,实现不重训练即可对齐特定用户观点。该方法成本低、效率高,同时保持模型原有推理能力。

原文摘要 · Abstract (English)

LLMs internally organize political ideology along low-dimensional structures that are partially, but not fully aligned with human ideological space. This misalignment is systematic, model specific, and measurable. We introduce a lightweight linear probe that both quantifies the misalignment and minimally corrects the output layer. This paper introduces a simple and efficient method for aligning models with specific user opinions. Instead of retraining the model, we calculated a bias score from its internal features and directly adjusted the final output probabilities. This solution is practical and low-cost and preserves the original reasoning power of the model.

意识形态对齐轻量修正社会媒体分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。