专家查文档不只看内容,还依赖社会语境和反复迭代。
Beyond Text: Characterizing Domain Expert Needs in Document Research
- 通过访谈16位专家,发现文档研究过程高度个性化且依赖社会背景。
- 现有NLP系统多聚焦文本内容,难匹配专家真实工作流程。
- 建议构建更可定制、可迭代、具社会感知的文档工具。
文档处理是知识工作的重要部分,如文献综述或法律判例审查。尽管近年来基于文本的NLP系统被宣传为能辅助甚至自动化此类任务,但其是否真正理解专家实际如何开展文档研究?本研究对两个领域共16位专家进行访谈,探究其文档研究流程,并与当前NLP系统能力对比。结果发现,专家的工作具有高度个性化、迭代性特征,且强烈依赖文档的社会语境,而不仅是内容本身。那些将文档视为独立对象而非仅文本容器的方法,更贴近专家关注点,但往往局限于特定研究圈层,外部难以使用。研究呼吁自然语言处理社区更重视文档在工具设计中的核心地位,推动开发可定制、可迭代、具社会感知的实用系统。
原文摘要 · Abstract (English)
Working with documents is a key part of almost any knowledge work, from contextualizing research in a literature review to reviewing legal precedent. Recently, as their capabilities have expanded, primarily text-based NLP systems have often been billed as able to assist or even automate this kind of work. But to what extent are these systems able to model these tasks as experts conceptualize and perform them now? In this study, we interview sixteen domain experts across two domains to understand their processes of document research, and compare it to the current state of NLP systems. We find that our participants processes are idiosyncratic, iterative, and rely extensively on the social context of a document in addition its content; existing approaches in NLP and adjacent fields that explicitly center the document as an object, rather than as merely a container for text, tend to better reflect our participants' priorities, though they are often less accessible outside their research communities. We call on the NLP community to more carefully consider the role of the document in building useful tools that are accessible, personalizable, iterative, and socially aware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。