Files changed (1) hide show
  1. README.md +215 -0
README.md ADDED
@@ -0,0 +1,215 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - es
4
+ license: mit
5
+ library_name: transformers
6
+ pipeline_tag: audio-classification
7
+ tags:
8
+ - emotion-recognition
9
+ - speech-emotion-recognition
10
+ - audio-classification
11
+ - speech-processing
12
+ - spanish
13
+ - affective-computing
14
+ - umuteam
15
+ datasets:
16
+ - NLP-UMUTeam/Spanish-MEACorpus-2023
17
+ metrics:
18
+ - accuracy
19
+ - f1
20
+
21
+ model-index:
22
+ - name: UMUTeam/w2v-bert-emotion-es
23
+ results:
24
+ - task:
25
+ type: audio-classification
26
+ name: Speech Emotion Recognition
27
+ dataset:
28
+ name: Spanish MEACorpus 2023
29
+ type: custom
30
+ metrics:
31
+ - type: accuracy
32
+ value: 88.1207
33
+ name: Accuracy
34
+ - type: weighted-f1
35
+ value: 88.1357
36
+ name: Weighted F1
37
+ - type: macro-f1
38
+ value: 84.4829
39
+ name: Macro F1
40
+ ---
41
+
42
+ # UMUTeam/w2v-bert-emotion-es
43
+
44
+ ## Model description
45
+
46
+ `UMUTeam/w2v-bert-emotion-es` is a Spanish speech emotion recognition model developed as part of **speech-emotion**, an open-source multilingual and multimodal toolkit for emotion recognition from speech, text, and multimodal inputs.
47
+
48
+ This model performs **emotion classification directly from Spanish speech audio**.
49
+
50
+ The model is based on the Wav2Vec2-BERT architecture and was fine-tuned for speech emotion recognition tasks in Spanish.
51
+
52
+ It is designed to operate as a standalone speech-only emotion recognition system or as part of the broader `speech-emotion` framework, where acoustic representations can be combined with textual representations for multimodal emotion recognition.
53
+
54
+ The model predicts one of the following emotion labels:
55
+
56
+ - `anger`
57
+ - `disgust`
58
+ - `fear`
59
+ - `joy`
60
+ - `neutral`
61
+ - `sadness`
62
+
63
+ ## Intended use
64
+
65
+ This model is intended for research and applied scenarios involving Spanish speech emotion recognition, such as:
66
+
67
+ - emotion analysis from speech recordings
68
+ - conversational speech analysis
69
+ - affective computing research
70
+ - human-computer interaction
71
+ - emotion-aware conversational agents
72
+ - integration into multimodal emotion recognition pipelines
73
+
74
+ It can be used directly with the Hugging Face `transformers` library or through the `speech-emotion` toolkit.
75
+
76
+ ## Out-of-scope use
77
+
78
+ This model should not be used as the sole basis for high-stakes decisions, including but not limited to:
79
+
80
+ - clinical diagnosis
81
+ - mental health assessment
82
+ - employment, legal, or educational decisions
83
+ - biometric profiling or surveillance
84
+ - automated decisions affecting individuals without human oversight
85
+
86
+ Emotion recognition is inherently uncertain and context-dependent. Predictions should be interpreted as model estimates, not as definitive assessments of a person's emotional state.
87
+
88
+ ## Training data
89
+
90
+ The model was trained on the Spanish portion of the datasets used in the `speech-emotion` project, primarily based on the **Spanish MEACorpus 2023** dataset.
91
+
92
+ Spanish MEACorpus 2023 is a multimodal speech-text emotion corpus for Spanish emotion analysis collected from natural environments. The dataset contains aligned speech and textual information for emotion recognition tasks.
93
+
94
+ The emotion labels were harmonized into the following six-class taxonomy:
95
+
96
+ - `anger`
97
+ - `disgust`
98
+ - `fear`
99
+ - `joy`
100
+ - `neutral`
101
+ - `sadness`
102
+
103
+ For the Spanish speech emotion recognition setup:
104
+
105
+ - Training samples: 3,692
106
+ - Validation samples: 410
107
+ - Test samples: 1,027
108
+
109
+ More details about the dataset and preprocessing pipeline are available in the project repository:
110
+
111
+ https://github.com/NLP-UMUTeam/umuteam-speech-emotion
112
+
113
+ ## Evaluation
114
+
115
+ The model was evaluated on the Spanish held-out test set used in the `speech-emotion` toolkit.
116
+
117
+ | Language | Mode | Accuracy | Weighted Precision | Weighted F1 | Macro F1 |
118
+ |---|---:|---:|---:|---:|---:|
119
+ | Spanish | Speech | 88.1207 | 88.3244 | 88.1357 | 84.4829 |
120
+
121
+ These results correspond to the speech-only Spanish configuration. In the full toolkit, multimodal configurations combining speech and text obtain even higher performance, showing the benefit of integrating acoustic and linguistic information.
122
+
123
+ ## How to use
124
+
125
+ ```python
126
+ from transformers import pipeline
127
+
128
+ classifier = pipeline(
129
+ "audio-classification",
130
+ model="UMUTeam/w2v-bert-emotion-es"
131
+ )
132
+
133
+ prediction = classifier("audio.wav")
134
+
135
+ print(prediction)
136
+ ```
137
+
138
+ You can also use this model through the `speech-emotion` toolkit:
139
+
140
+ ```bash
141
+ pip install speech-emotion
142
+ ```
143
+
144
+ ```python
145
+ from speech_emotion import predict_emotion
146
+
147
+ emotion = predict_emotion(
148
+ audio_path="audio.wav",
149
+ language="es",
150
+ mode="audio",
151
+ model_config_path="model.json"
152
+ )
153
+
154
+ print("Detected emotion:", emotion)
155
+ ```
156
+
157
+ Repository:
158
+
159
+ https://github.com/NLP-UMUTeam/umuteam-speech-emotion
160
+
161
+ ## Limitations
162
+
163
+ - The model is designed for Spanish speech and may not perform reliably on other languages.
164
+ - It predicts a single label from a fixed set of six emotions.
165
+ - Emotion expression is subjective and highly context-dependent.
166
+ - Performance may decrease with noisy audio, overlapping speakers, low-quality recordings, strong accents, or domain shifts.
167
+ - Speech-only emotion recognition may miss relevant contextual or visual information that could improve emotion interpretation.
168
+
169
+ ## Bias and ethical considerations
170
+
171
+ Emotion recognition systems may reflect biases present in their training data, including differences related to accents, speaking styles, demographics, recording conditions, or annotation subjectivity.
172
+
173
+ Users should avoid interpreting predictions as objective truths about a person's internal emotional state. The model should be used with transparency, appropriate consent, and human oversight, especially in sensitive contexts.
174
+
175
+ ## Citation
176
+
177
+ If you use this model in your research, please cite the following works:
178
+
179
+ ### speech-emotion toolkit
180
+
181
+ ```bibtex
182
+ @article{PAN2026102677,
183
+ title = {speech-emotion: A multilingual and multimodal toolkit for emotion recognition from speech},
184
+ journal = {SoftwareX},
185
+ volume = {34},
186
+ pages = {102677},
187
+ year = {2026},
188
+ issn = {2352-7110},
189
+ doi = {https://doi.org/10.1016/j.softx.2026.102677},
190
+ url = {https://www.sciencedirect.com/science/article/pii/S235271102600169X},
191
+ author = {Ronghao Pan and Tomás Bernal-Beltrán and José Antonio García-Díaz and Rafael Valencia-García},
192
+ }
193
+ ```
194
+
195
+ ### Spanish MEACorpus 2023
196
+
197
+ ```bibtex
198
+ @article{PAN2024103856,
199
+ title = {Spanish MEACorpus 2023: A multimodal speech–text corpus for emotion analysis in Spanish from natural environments},
200
+ journal = {Computer Standards & Interfaces},
201
+ volume = {90},
202
+ pages = {103856},
203
+ year = {2024},
204
+ issn = {0920-5489},
205
+ doi = {https://doi.org/10.1016/j.csi.2024.103856},
206
+ url = {https://www.sciencedirect.com/science/article/pii/S0920548924000254},
207
+ author = {Ronghao Pan and José Antonio García-Díaz and Miguel Ángel Rodríguez-García and Rafael Valencia-García},
208
+ }
209
+ ```
210
+
211
+ ## Acknowledgments
212
+
213
+ This work is part of the research project LaTe4PoliticES (PID2022-138099OB-I00), funded by MICIU/AEI/10.13039/501100011033 and the European Regional Development Fund (ERDF/EU - FEDER/UE), “A way of making Europe”.
214
+
215
+ Mr. Tomás Bernal-Beltrán is supported by the University of Murcia through the predoctoral programme.