Semi-supervised cross-lingual speech emotion recognition

Agarla, Mirko; Bianco, Simone; Celona, Luigi; Napoletano, Paolo; Petrovsky, Alexey; Piccoli, Flavio; Schettini, Raimondo; Shanin, Ivan

doi:10.1016/j.eswa.2023.121368

Computer Science > Sound

arXiv:2207.06767 (cs)

[Submitted on 14 Jul 2022 (v1), last revised 17 Jul 2023 (this version, v2)]

Title:Semi-supervised cross-lingual speech emotion recognition

Authors:Mirko Agarla, Simone Bianco, Luigi Celona, Paolo Napoletano, Alexey Petrovsky, Flavio Piccoli, Raimondo Schettini, Ivan Shanin

View PDF

Abstract:Performance in Speech Emotion Recognition (SER) on a single language has increased greatly in the last few years thanks to the use of deep learning techniques. However, cross-lingual SER remains a challenge in real-world applications due to two main factors: the first is the big gap among the source and the target domain distributions; the second factor is the major availability of unlabeled utterances in contrast to the labeled ones for the new language. Taking into account previous aspects, we propose a Semi-Supervised Learning (SSL) method for cross-lingual emotion recognition when only few labeled examples in the target domain (i.e. the new language) are available. Our method is based on a Transformer and it adapts to the new domain by exploiting a pseudo-labeling strategy on the unlabeled utterances. In particular, the use of a hard and soft pseudo-labels approach is investigated. We thoroughly evaluate the performance of the proposed method in a speaker-independent setup on both the source and the new language and show its robustness across five languages belonging to different linguistic strains. The experimental findings indicate that the unweighted accuracy is increased by an average of 40% compared to state-of-the-art methods.

Subjects:	Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2207.06767 [cs.SD]
	(or arXiv:2207.06767v2 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2207.06767
Journal reference:	Elsevier Expert Systems with Applications, 237 (2024), 121368
Related DOI:	https://doi.org/10.1016/j.eswa.2023.121368

Submission history

From: Luigi Celona [view email]
[v1] Thu, 14 Jul 2022 09:24:55 UTC (283 KB)
[v2] Mon, 17 Jul 2023 06:11:59 UTC (2,677 KB)

Computer Science > Sound

Title:Semi-supervised cross-lingual speech emotion recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Semi-supervised cross-lingual speech emotion recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators