Skip to main content
Dryad

KeyBench: A resource for extracting Keywords from scientific papers

Data files

Jul 13, 2026 version files 91.27 MB

Click names to download individual files

Abstract

This resource is intended to encourage the community to build tools to extract keywords. Many papers on ArXiv list keywords on the first page, typically near the abstract. The proposed resource scrapes these keywords, and uses them as the gold standard for a prediction task.  An evaluation compares three baselines for predicting the gold standard: (1) VitaLity, (2) topics from Semantic Scholar and (3) a chatbot. We challenge the community to beat these baselines. To make it easier for the community to build such systems, the resource includes a number of fields for 117,877 papers from Semantic Scholar: (a) title, (b) abstract, (c) authors, (d) citing sentences, and (e) ids to make it easier to join with data in resources such as ArXiv, Semantic Scholar, ACL Anthology, PubMed, etc. We hope the community will show that citing sentences are useful because of the wisdom of the crowd (good keywords are likely to be used by many authors).