Article Info

Quantifying Semantic Shift Visually on a Malay Domain Specific Corpus Using Temporal Word Embedding Approach

Sabrina Tiun, Saidah Saad, Nor Fariza Mohd Noor, Azhar Jalaludin, Anis Nadiah Che Abdul Rahman
dx.doi.org/10.17576/apjitm-2020-0902-01

Abstract

In this study, we propose an alternative approach to analyzing a domain-specific time series corpus for detecting word evolution. The method trains a target corpus in time series into a temporal word embedding (TWE) model. The advantage of TWE is that one can see how the meaning of a word changes over time. We have chosen the TWEC approach to model a Malay domain-specific time-series corpus, the Malaysian Hansard Corpus (MHC), to a TWE model and called the model as MHC-TWEC. Two primary analyses, i.e., self-similarity analysis and user-defined method analysis, were performed to validate the effectiveness of the MHC-TWEC model in quantifying semantic shift on MHC visually. From those analyses, we visually find out that the TWE model can capture the semantic shift in the temporal corpus (the MHC).

keyword

temporal word embedding, temporal corpus, Malaysian Hansard Corpus

Area

Knowledge Technology