Unsupervised Separation of Transliterable and Native Words for Malayalam

Research output: Chapter in Book/Report/Conference proceedingConference contribution

164 Downloads (Pure)

Abstract

Differentiating intrinsic language words from transliterable words is a key step aiding text processing tasks involving different natural languages. We consider the problem of unsupervised separation of transliterable words from native words for text in Malayalam language. Outlining a key observation on the diversity of characters beyond the word stem, we develop an optimization method to score words based on their nativeness. Our method relies on the usage of probability distributions over character n-grams that are refined in step with the nativeness scorings in an iterative optimization formulation. Using an empirical evaluation, we illustrate that our method, DTIM, provides significant improvements in nativeness scoring for Malayalam, establishing DTIM as the preferred method for the task.
Original languageEnglish
Title of host publicationProceedings of the 14th International Conference on Natural Language Processing (ICON 2017)
Pages155-164
Number of pages10
Publication statusPublished - 21 Dec 2017
EventICON 2017 - Kolkata, Kolkata, India
Duration: 18 Dec 201721 Dec 2017
https://ltrc.iiit.ac.in/icon2017/

Conference

ConferenceICON 2017
Country/TerritoryIndia
CityKolkata
Period18/12/201721/12/2017
Internet address

Fingerprint

Dive into the research topics of 'Unsupervised Separation of Transliterable and Native Words for Malayalam'. Together they form a unique fingerprint.

Cite this