Abstract
Automatic subject prediction is a desirable feature for modern digital library systems, as manual indexing can no longer cope with the rapid growth of digital collections. It is also desirable to be able to identify a small set of entities (e.g., authors, citations, bibliographic records) which are most relevant to a query. This gets more difficult when the amount of data increases dramatically. Data sparsity and model scalability are the major challenges to solving this type of extreme multi-label classification problem automatically. In this paper, we propose to address this problem in two steps: we first embed different types of entities into the same semantic space, where similarity could be computed easily; second, we propose a novel non-parametric method to identify the most relevant entities in addition to direct semantic similarities. We show how effectively this approach predicts even very specialised subjects, which are associated with few documents in the training set and are more problematic for a classifier.
| Original language | English |
|---|---|
| Pages (from-to) | 364-370 |
| Number of pages | 7 |
| Journal | Knowledge Organization |
| Volume | 46 |
| Issue number | 5 |
| DOIs | |
| Publication status | Published - 2019 |
| Externally published | Yes |
Keywords
- Documents
- Embedding
- Entities
- Subjects
Fingerprint
Dive into the research topics of 'Embed First, Then Predict'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver