An effective cybernated word embedding system for analysis and language identification in code-mixed social media text

Abstract Social media texts like tweets and blogs are collaboratively created by human interaction. Fast change in trends leads to topic drift in the social media text. This drift is usually associated with words and hashtags. However, geotags play an important part in determining topic distribution with location context. Rate of change in the distribution of words, hashtags and geotags cannot be considered as uniform and must be handled accordingly. This paper builds a topic model that associates topic with a mixture of distributions of words, hashtags and geotags. Stochastic gradient Langevin dynamic model with varying mini-batch sizes is used to capture the changes due to the asynchronous distribution of words and tags. Topical word embedding with co-occurrence and location contexts are specified as hashtag context vector and geotag context vector respectively. These two vectors are jointly learned to yield topical word embedding vectors related to tags context. Topical word embeddings over time conditioned on hashtags and geotags predict, location-based topical variations effectively. When evaluated with Chennai and UK geolocated Twitter data, the proposed joint topical word embedding model enhanced by the social tags context, outperforms other methods.

Download Full-text

Joint Topical Word Embedding for Detecting Drift in Social Media Text

10.21203/rs.3.rs-90835/v1 ◽

2020 ◽

Author(s):

VIJAYARANI J ◽

Geetha T.V.

Keyword(s):

Social Media ◽

Topic Model ◽

Rate Of Change ◽

Word Embedding ◽

Human Interaction ◽

Langevin Dynamic ◽

Context Vector ◽

The Social ◽

Topic Distribution ◽

Social Media Text

Abstract Social media texts like tweets and blogs are collaboratively created by human interaction. Fast change in trends leads to topic drift in the social media text. This drift is usually associated with words and hashtags. However, geotags play an important part in determining topic distribution with location context. Rate of change in the distribution of words, hashtags and geotags cannot be considered as uniform and must be handled accordingly. This paper builds a topic model that associates topic with a mixture of distributions of words, hashtags and geotags. Stochastic gradient Langevin dynamic model with varying mini-batch sizes is used to capture the changes due to the asynchronous distribution of words and tags. Topical word embedding with co-occurrence and location contexts are specified as hashtag context vector and geotag context vector respectively. These two vectors are jointly learned to yield topical word embedding vectors related to tags context. Topical word embeddings over time conditioned on hashtags and geotags predict, location-based topical variations effectively. When evaluated with Chennai and UK geolocated Twitter data, the proposed joint topical word embedding model enhanced by the social tags context, outperforms other methods.

Download Full-text

An Automatic Language Identification System for Code-Mixed English-Kannada Social Media Text

2017 2nd International Conference on Computational Systems and Information Technology for Sustainable Solution (CSITSS) ◽

10.1109/csitss.2017.8447784 ◽

2017 ◽

Author(s):

B S Sowmya Lakshmi ◽

B R Shambhavi

Keyword(s):

Social Media ◽

Language Identification ◽

Identification System ◽

Social Media Text

Download Full-text

An Effective Bi-LSTM Word Embedding System for Analysis and Identification of Language in Code-Mixed social Media Text in English and Roman Hindi

Computación y Sistemas ◽

10.13053/cys-24-4-3151 ◽

2020 ◽

Vol 24 (4) ◽

Author(s):

Shashi Shekhar ◽

Dilip Kumar Sharma ◽

M.M. Sufyan Beg

Keyword(s):

Social Media ◽

Word Embedding ◽

Social Media Text

Download Full-text

Character Embedding for Language Identification in Hindi-English Code-mixed Social Media Text

Computación y Sistemas ◽

10.13053/cys-22-1-2775 ◽

2018 ◽

Vol 22 (1) ◽

Cited By ~ 5

Author(s):

P. V. Veena ◽

M. Anand Kumar ◽

K. P. Soman

Keyword(s):

Social Media ◽

Language Identification ◽

Social Media Text

Download Full-text

Language Identification and Analysis of Code-Switched Social Media Text

10.18653/v1/w18-3206 ◽

2018 ◽

Cited By ~ 4

Author(s):

Deepthi Mave ◽

Suraj Maharjan ◽

Thamar Solorio

Keyword(s):

Social Media ◽

Language Identification ◽

Social Media Text

Download Full-text

Deep Learning-Based Language Identification in English-Hindi-Bengali Code-Mixed Social Media Corpora

Journal of Intelligent Systems ◽

10.1515/jisys-2017-0440 ◽

2019 ◽

Vol 28 (3) ◽

pp. 399-408 ◽

Cited By ~ 3

Author(s):

Anupam Jamatia ◽

Amitava Das ◽

Björn Gambäck

Keyword(s):

Social Media ◽

Deep Learning ◽

Conditional Random Fields ◽

Short Term Memory ◽

Language Identification ◽

Code Mixing ◽

Word Level ◽

Social Media Text ◽

Feature Based ◽

Bidirectional Lstm

Abstract This article addresses language identification at the word level in Indian social media corpora taken from Facebook, Twitter and WhatsApp posts that exhibit code-mixing between English-Hindi, English-Bengali, as well as a blend of both language pairs. Code-mixing is a fusion of multiple languages previously mainly associated with spoken language, but which social media users also deploy when communicating in ways that tend to be rather casual. The coarse nature of code-mixed social media text makes language identification challenging. Here, the performance of deep learning on this task is compared to feature-based learning, with two Recursive Neural Network techniques, Long Short Term Memory (LSTM) and bidirectional LSTM, being contrasted to a Conditional Random Fields (CRF) classifier. The results show the deep learners outscoring the CRF, with the bidirectional LSTM demonstrating the best language identification performance.

Download Full-text