Publication

C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content

Journal Contribution - Journal Article

We study the problem of extracting cross-lingual topics from non-parallel multilingual text datasets with partially overlapping thematic content (e.g., aligned Wikipedia articles in two different languages). To this end, we develop a new bilingual probabilistic topic model called comparable bilingual latent Dirichlet allocation (C-BiLDA), which is able to deal with such comparable data, and, unlike the standard bilingual LDA model (BiLDA), does not assume the availability of document pairs with identical topic distributions. We present a full overview of C-BiLDA, and show its utility in the task of cross-lingual knowledge transfer for multi-class document classification on two benchmarking datasets for three language pairs. The proposed model outperforms the baseline LDA model, as well as the standard BiLDA model and two standard low-rank approximation methods (CL-LSI and CL-KCCA) used in previous work on this task.

Journal: Data Mining and Knowledge Discovery

ISSN: 1384-5810

Issue: 5

Volume: 30

Pages: 1299 - 1323

Publication year:2016

Institutional Repository URL: https://lirias.kuleuven.be/931579
DOI: https://doi.org/10.1007/s10618-015-0442-x
WoS Id: 000382010500013

BOF-keylabel:yes

IOF-keylabel:yes

BOF-publication weight:1

CSS-citation score:1

Authors from:Higher Education

Accessibility:Open

Publication

C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content

Journal Contribution - Journal Article

Authors/publisher

Research units