ProQuest
Abstract/Details

Context, Bias, and Dynamics in Speech Emotion Recognition

Martinez-Lucas, Luz M.   The University of Texas at Dallas ProQuest Dissertations & Theses,  2026. 32450640.

Abstract (summary)

Affective computing relies on high-quality emotional annotations to analyze, recognize, and model human emotions. In speech emotion recognition (SER), these annotations are commonly collected at the sentence-level. However, emotional perception is inherently dynamic and context-dependent. Annotation processes are susceptible to cognitive biases such as the affective priming effect, in which previously perceived emotions influence subsequent judgments. This dissertation first investigates how affective priming impacts emotional annotations and examines its implications for SER systems. We show that ratings are biased toward previously perceived emotional extremes and demonstrate that SER models trained on the most biased samples achieve higher performance and lower prediction uncertainty. To support the study of context-dependent and dynamic emotions, we introduce the MSP-Conversation corpus, a large-scale dataset containing over 70 hours of conversational speech with time-continuous annotations of valence, arousal, and dominance, along with detailed speaker diarization. The corpus overlaps with the MSP-Podcast dataset, enabling direct comparisons between in-context time-continuous annotations and out-of-context sentence-level annotations. Using shared recordings, we analyze the similarity, agreement, and SER performance of sentence-level labels derived from time-continuous annotations versus labels collected directly at the sentence level.

Finally, this work advances dynamic speech emotion recognition (DSER) by proposing context-aware modeling approaches. We present a conditional neural process framework that learns priors over sparse observations of the emotional signal and significantly outperforms a BiLSTM baseline in predicting emotional traces. We further introduce a DSER model that incorporates the same observations using attention mechanisms, achieving higher concordance correlation coefficients. This formulation demonstrates the possible use of our final model in human-inthe-loop DSER, where sparse and impactful human feedback guides the emotional predictions. Together, these contributions highlight the importance of contextual bias, dynamic annotation, and context-aware modeling for robust and realistic SER.

Indexing (details)


Business indexing term
Subject
Electrical engineering;
Computer science;
Artificial intelligence
Classification
0800: Artificial intelligence
0544: Electrical engineering
0984: Computer science
Identifier / keyword
Affective computing; Computational paralinguistics; Speech emotion recognition; Emotional annotation; Machine learning; Speech technology
Title
Context, Bias, and Dynamics in Speech Emotion Recognition
Author
Martinez-Lucas, Luz M.  VIAFID ORCID Logo 
Number of pages
202
Publication year
2026
Degree date
2026
School code
0382
Source
DAI-B 87/12(E), Dissertation Abstracts International
ISBN
9798252444468
Advisor
Busso, Carlos; Hansen, John H. L.
Committee member
Sisman, Berrak; Panahi, Issa M. S.; Nosratinia, Aria
University/institution
The University of Texas at Dallas
Department
Electrical Engineering
University location
United States -- Texas
Degree
Ph.D.
Funding Agency and Grant Number
Funding Agency and Grant Number
Funding AgencyGrant Number
U.S. National Science FoundationCNS-1823166;CNS-2016719;IIS-1453781
Source type
Dissertation or Thesis
Language
English
Document type
Dissertation/Thesis
Dissertation/thesis number
32450640
ProQuest document ID
3355061848
Copyright
Database copyright ProQuest LLC; ProQuest does not claim copyright in the individual underlying works.
Document URL
https://www.proquest.com/pqdtglobal/docview/3355061848/7251C2AF461F413EPQ/4