< Back to previous page

Publication

Database Principles and Challenges in Text Analysis

Journal Contribution - Journal Article

A common conceptual view of text analysis is that of a two-step process, where we first extract relations from text documents and then apply a relational query over the result. Hence, text analysis shares technical challenges with, and can draw ideas from, relational databases. A framework that formally instantiates this connection is that of the document spanners. In this article, we review recent advances in various research efforts that adapt fundamental database concepts to text analysis through the lens of document spanners. Among others, we discuss aspects of query evaluation, aggregate queries, provenance, and distributed query planning.
Journal: SIGMOD RECORD
ISSN: 0163-5808
Issue: 2
Volume: 50
Pages: 6 - 17
Publication year:2021
BOF-keylabel:yes
IOF-keylabel:yes
BOF-publication weight:0.1
Authors:International
Authors from:Higher Education
Accessibility:Closed