QUANTIFYING SEMANTIC SHIFT VISUALLY ON A MALAY DOMAIN SPECIFIC CORPUS USING TEMPORAL WORD EMBEDDING APPROACH

In this study, we propose an alternative approach to analyzing a domain-specific time series corpus for detecting word evolution. The method trains a target corpus in time series into a temporal word embedding (TWE) model. The advantage of TWE is that one can see how the meaning of a word changes ov...

Full description

Bibliographic Details
Main Authors: Sabrina Tiun, Saidah Saad, Nor Fariza Mohd Noor, Azhar Jalaludin, Anis Nadiah Che Abdul Rahman
Format: Article
Language:English
Published: UKM Press 2020-12-01
Series:Asia-Pacific Journal of Information Technology and Multimedia
Subjects:
Online Access:https://www.ukm.my/apjitm/view.php?id=197
Description
Summary:In this study, we propose an alternative approach to analyzing a domain-specific time series corpus for detecting word evolution. The method trains a target corpus in time series into a temporal word embedding (TWE) model. The advantage of TWE is that one can see how the meaning of a word changes over time. We have chosen the TWEC approach to model a Malay domain-specific time-series corpus, the Malaysian Hansard Corpus (MHC), to a TWE model and called the model as MHC-TWEC. Two primary analyses, i.e., self-similarity analysis and user-defined method analysis, were performed to validate the effectiveness of the MHC-TWEC model in quantifying semantic shift on MHC visually. From those analyses, we visually find out that the TWE model can capture the semantic shift in the temporal corpus (the MHC).
ISSN:2289-2192