Context, Bias, and Dynamics in Speech Emotion Recognition
Abstract (summary)
Affective computing relies on high-quality emotional annotations to analyze, recognize, and model human emotions. In speech emotion recognition (SER), these annotations are commonly collected at the sentence-level. However, emotional perception is inherently dynamic and context-dependent. Annotation processes are susceptible to cognitive biases such as the affective priming effect, in which previously perceived emotions influence subsequent judgments. This dissertation first investigates how affective priming impacts emotional annotations and examines its implications for SER systems. We show that ratings are biased toward previously perceived emotional extremes and demonstrate that SER models trained on the most biased samples achieve higher performance and lower prediction uncertainty. To support the study of context-dependent and dynamic emotions, we introduce the MSP-Conversation corpus, a large-scale dataset containing over 70 hours of conversational speech with time-continuous annotations of valence, arousal, and dominance, along with detailed speaker diarization. The corpus overlaps with the MSP-Podcast dataset, enabling direct comparisons between in-context time-continuous annotations and out-of-context sentence-level annotations. Using shared recordings, we analyze the similarity, agreement, and SER performance of sentence-level labels derived from time-continuous annotations versus labels collected directly at the sentence level.
Finally, this work advances dynamic speech emotion recognition (DSER) by proposing context-aware modeling approaches. We present a conditional neural process framework that learns priors over sparse observations of the emotional signal and significantly outperforms a BiLSTM baseline in predicting emotional traces. We further introduce a DSER model that incorporates the same observations using attention mechanisms, achieving higher concordance correlation coefficients. This formulation demonstrates the possible use of our final model in human-inthe-loop DSER, where sparse and impactful human feedback guides the emotional predictions. Together, these contributions highlight the importance of contextual bias, dynamic annotation, and context-aware modeling for robust and realistic SER.
Indexing (details)
Computer science;
Artificial intelligence
0544: Electrical engineering
0984: Computer science
| Funding Agency | Grant Number |
|---|---|
| U.S. National Science Foundation | CNS-1823166;CNS-2016719;IIS-1453781 |
