WEXEA: Wikipedia EXhaustive Entity Annotation

Abstract

In this paper, we are discussing an approach in order to create a text corpus based on Wikipedia with exhaustive annotations of entity mentions. Editors on Wikipedia are only expected to add hyperlinks in order to help the reader to understand the content, but are discouraged to add links that do not add any benefit for understanding an article. Therefore, many mentions of popular entities (such as countries or popular events in history), previously linked articles as well as the article entity itself, are not linked. This results in a huge potential for additional annotations that can be used for downstream NLP tasks, such as Relation Extraction. We show that our annotations are useful for creating distantly supervised datasets for this task. Furthermore, we publish all code necessary to derive a corpus from a raw Wikipedia dump, so that it can be reproduced by everyone.

WEXEA: Wikipedia EXhaustive Entity Annotation

Abstract

Latest Research Papers

Basic and Depression Specific Emotions Identification in Tweets: Multi-label Classification Experiments

Weakly-Supervised Questions for Zero-Shot Relation Extraction

Updating displayed data visualizations according to identified conversation centers in natural language commands

Let us help you

Connect with the community

Explore training and advanced education

Harness the potential of artificial intelligence

Connect with the community

Explore training and advanced education

Harness the potential of artificial intelligence