小模型通过页面分词更懂用户浏览偏好,效果优于大模型。
Subjective Behaviors and Preferences in LLM: Language of Browsing
- 用页面级分词训练小型语言模型,模拟用户浏览‘语言’。
- 小模型在行为预测上超越大型预训练模型,准确率更高。
- 针对多样化用户设计异质参数,提升个性化适配与生成一致性。
大型语言模型(LLM)虽被广泛认为能适应多样化的用户行为与偏好,但用户在浏览网页或应用时的主观性行为常具独特性。这些行为日志构成类似用户自建的‘浏览语言’,缺乏自然语言的结构与语法。本文提出三个问题:(i) 小型模型能否比大型模型更好地表征这种‘浏览语言’?(ii) 单一参数集的模型是否足以捕捉多样的用户偏好?(iii) 高平均性能且低方差的模型能否实现更好的用户层面对齐?为此,我们引入面向主观行为的异质性感知训练方法 HeTLM。实验发现:(i) 使用页面级分词的小型模型优于大型预训练或微调模型;(ii) 采用异质聚类专属参数的 HeTLM 模型,在参数量相当下优于单一模型;(iii) 模型生成表现具有更高的均值与更低的方差,表明其在用户层级的对齐能力更强。
原文摘要 · Abstract (English)
A Large Language Model (LLM) offers versatility across domains and tasks, purportedly benefiting users with a wide variety of behaviors and preferences. We question this perception about an LLM when users have inherently subjective behaviors and preferences, as seen in their ubiquitous and idiosyncratic browsing of websites or apps. The sequential behavior logs of pages, thus generated, form something akin to each user's self-constructed "language", albeit without the structure and grammar imbued in natural languages. We ask: (i) Can a small LM represent the "language of browsing" better than a large LM? (ii) Can an LM with a single set of parameters (or, single LM) adequately capture myriad users' heterogeneous, subjective behaviors and preferences? (iii) Can a single LM with high average performance, yield low variance in performance to make alignment good at user level? We introduce clusterwise LM training, HeTLM (Heterogeneity aware Training of Language Model), appropriate for subjective behaviors. We find that (i) a small LM trained using a page-level tokenizer outperforms large pretrained or finetuned LMs; (ii) HeTLM with heterogeneous cluster specific set of parameters outperforms a single LM of the same family, controlling for the number of parameters; and (iii) a higher mean and a lower variance in generation ensues, implying improved alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。