Data from: Parser extraction of triples in unstructured text

D'Souza, Shaun1

Published Apr 26, 2019 on Dryad. https://doi.org/10.5061/dryad.s7j17qp

Data files

Apr 26, 2019 version files 11.88 KB

Parser Extraction of Triples in Unstructured Text.zip

11.88 KB

Abstract

The web contains vast repositories of unstructured text. We investigate the opportunity for building a knowledge graph from these text sources. We generate a set of triples which can be used in knowledge gathering and integration. We define the architecture of a language compiler for processing subject-predicate-object triples using the OpenNLP parser. We implement a depth-first search traversal on the POS tagged syntactic tree appending predicate and object information. A parser enables higher precision and higher recall extractions of syntactic relationships across conjunction boundaries. We are able to extract 2-2.5 times the correct extractions of ReVerb. The extractions are used in a variety of semantic web applications and question answering. We verify extraction of 50,000 triples on the ClueWeb dataset.

Data from: Parser extraction of triples in unstructured text

Data files

Abstract

Usage notes

Parser Extraction of Triples in Unstructured Text

Works referencing this dataset