Speech Emotion Recognition with Whisper: Classifier Head vs Vocabulary Token
audio
speech
architecture
Two designs over CREMA-D, three seeds each, and what changes when the label is a class index instead of a word the decoder has to say.
Author
Rosh Beed
Published
July 9, 2026
Whisper transcribes. It does not tell you how something was said, and how is often the point: the same sentence read angrily and read sadly is one transcript and two different messages.
The obvious route is a classifier.
Take the encoder
Average its outputs into one vector
Put a single linear layer on top, one score per emotion
Train it so the right emotion scores highest
Three lines, and it works.
I built it the other way, because of something Whisper already does.
Whisper’s Token Prefix
A Whisper transcript does not start with words. It starts with a structured token prefix, and that prefix is what lets one model do transcription, translation and language identification without being three models.
Those are not a separate mechanism bolted on. They are ordinary vocabulary entries, predicted by the same softmax and trained by the same cross-entropy as any word. Look at the ids above:
<|startoftranscript|><|en|><|transcribe|><|notimestamps|>Hello, my name is Bes.<|endoftext|>
[50258, 50259, 50359, 50363, 15947, 11, 452, 1315, 307, 8190, 13, 50257]
The first four are the language and the task. Nothing distinguishes them from 15947 except what the model learned to do when it sees them.
So: is emotion a different kind of thing from language and task, or the same kind of thing? It is hard to argue it is different. Adding <emotion_happy> and its siblings to that vocabulary lets the model say what it heard through machinery it already has.
Show the code
import reimport numpy as npimport torchimport torch.nn as nnfrom huggingface_hub import hf_hub_downloadfrom transformers import WhisperForConditionalGeneration, WhisperProcessorDEVICE = torch.device("mps"if torch.backends.mps.is_available() else"cpu")MODEL ="openai/whisper-small"BATCH, LR, EPOCHS, SEEDS =4, 1e-5, 3, (0, 1, 2)CREMAD ="3a5d5fa8a9945e122258fb6c52c243126186c227"data = np.load(hf_hub_download("roshbeed/ai-residency-blog-data", "speech/cremad.npz", repo_type="dataset", revision=CREMAD), allow_pickle=True)offsets, labels = data["offsets"], list(data["labels"])splits, EMOTIONS =list(data["splits"]), [str(e) for e in data["emotions"]]# Read the audio out once. np.load on a compressed .npz is lazy, so indexing# data["audio"] inside the loop decompresses the whole array on every iteration.# Slices of the one array, not copies of it. The conversion to float happens a# batch at a time, where it is needed.samples = data["audio"]clips = [samples[offsets[i]:offsets[i +1]] for i inrange(len(labels))]# CREMA-D is twelve fixed sentences said by ninety-one actors. The sentence is in# the filename, which is how the transcript is recovered.SENTENCES = {"IEO": "It's eleven o'clock", "TIE": "That is exactly what happened","IOM": "I'm on my way to the meeting", "IWW": "I wonder what this is about","TAI": "The airplane is almost full", "MTI": "Maybe tomorrow it will be cold","IWL": "I would like a new alarm clock", "ITH": "I think I have a doctor's appointment","DFA": "Don't forget a jacket", "ITS": "I think I've seen this before","TSI": "The surface is slick", "WSI": "We'll stop in a couple of minutes",}FILES ="180ad0a0636f15b42fb349d77a00758a3343d847"names = np.load(hf_hub_download("roshbeed/ai-residency-blog-data","speech/cremad-files.npz", repo_type="dataset", revision=FILES), allow_pickle=True)["files"]transcripts = [SENTENCES[str(f).rsplit("/", 1)[-1].split("_")[1]] for f in names]print(f"{len(clips):,} clips from CREMA-D, {len(EMOTIONS)} emotions: "f"{', '.join(EMOTIONS)}")print(f"{splits.count('train'):,} train, {splits.count('test'):,} test")print(f"guessing the emotion: {1/len(EMOTIONS):.3f}")
7,442 clips from CREMA-D, 6 emotions: anger, disgust, fear, happy, neutral, sad
5,209 train, 1,117 test
guessing the emotion: 0.167
Two designs, one model. Both start from the same released whisper-small weights; the only difference is where the emotion comes out.
The head design hangs a classifier off the encoder. The token design adds six entries to the vocabulary and makes the decoder emit one before it writes the transcript, so the emotion is predicted by the same softmax as every other token and trained by the same loss.
Show the code
processor = WhisperProcessor.from_pretrained(MODEL)EMOTION_TOKENS = [f"<|{e}|>"for e in EMOTIONS]def build(design, seed):"""One model, either with a classifier head or with six new vocabulary rows.""" torch.manual_seed(seed) tokeniser = WhisperProcessor.from_pretrained(MODEL).tokenizer# `generate` forces <|en|><|transcribe|><|notimestamps|> after the start token.# The tokenizer emits neither the language nor the task by default, so a target# built without them trains the decoder against a prefix that generation never# produces. Transcription survives it; a single token at a fixed position does# not, because the position it was trained at is one generation never reaches. tokeniser.set_prefix_tokens(language="en", task="transcribe") model = WhisperForConditionalGeneration.from_pretrained(MODEL) head =Noneif design =="token": tokeniser.add_tokens(EMOTION_TOKENS, special_tokens=True) model.resize_token_embeddings(len(tokeniser))else: head = nn.Linear(model.config.d_model, len(EMOTIONS))return model.to(DEVICE), (head.to(DEVICE) if head elseNone), tokeniserSOT = processor.tokenizer.convert_tokens_to_ids("<|startoftranscript|>")def targets(tokeniser, indices, design):"""Whisper prepends its own start token, so the tokenizer's is dropped."""if design =="token": text = [f"<|{labels[i]}|>{transcripts[i]}"for i in indices]else: text = [transcripts[i] for i in indices] ids = tokeniser(text, return_tensors="pt", padding=True).input_idsassert ids[0, 0].item() == SOTreturn ids[:, 1:].to(DEVICE)def features_for(indices): audio = [clips[i].astype(np.float32) /32768.0for i in indices]return processor(audio, sampling_rate=16000, return_tensors="pt").input_features.to(DEVICE)train_rows = [i for i, s inenumerate(splits) if s =="train"]test_rows = [i for i, s inenumerate(splits) if s =="test"]print(f"{len(train_rows):,} training clips, {len(test_rows):,} held out")print(f"the token design adds {len(EMOTION_TOKENS)} rows to a "f"{len(processor.tokenizer):,}-entry vocabulary")
5,209 training clips, 1,117 held out
the token design adds 6 rows to a 51,865-entry vocabulary
Training differs only in where the emotion sits. The token design puts it in the sequence and uses one cross-entropy over everything. The head design takes it out of the sequence and adds a second loss on the classifier.
Two numbers come back for each: how often the emotion is right, and how much of the transcript survives. The whole question is what the second costs.
Show the code
import gcimport timedef normalise(text):return re.sub(r"[^a-z' ]", "", text.lower()).split()def edits(reference, hypothesis): previous =list(range(len(hypothesis) +1))for i, want inenumerate(reference, 1): current = [i]for j, got inenumerate(hypothesis, 1): current.append(min(previous[j] +1, current[j -1] +1, previous[j -1] + (want != got))) previous = currentreturn previous[-1]def run(design, seed): model, head, tokeniser = build(design, seed) parameters =list(model.parameters()) + (list(head.parameters()) if head else []) optimiser = torch.optim.AdamW(parameters, lr=LR) generator = torch.Generator().manual_seed(seed)for _ inrange(EPOCHS): order = torch.randperm(len(train_rows), generator=generator).tolist()for start inrange(0, len(order) - BATCH +1, BATCH): batch = [train_rows[i] for i in order[start:start + BATCH]] features = features_for(batch) out = model(input_features=features, labels=targets(tokeniser, batch, design), output_hidden_states=(head isnotNone)) loss = out.lossif head isnotNone: pooled = out.encoder_last_hidden_state.mean(1) wanted = torch.tensor([EMOTIONS.index(labels[i]) for i in batch], device=DEVICE) loss = loss + nn.functional.cross_entropy(head(pooled), wanted) optimiser.zero_grad() loss.backward() optimiser.step() model.eval() right = errors = words =0for start inrange(0, len(test_rows), 16): batch = test_rows[start:start +16] features = features_for(batch)with torch.no_grad(): tokens = model.generate(features, max_new_tokens=48, language="en", task="transcribe")if head isnotNone: pooled = model.model.encoder(features).last_hidden_state.mean(1) guessed = head(pooled).argmax(-1).tolist() written = tokeniser.batch_decode(tokens, skip_special_tokens=False)for position, i inenumerate(batch):if head isnotNone: right += EMOTIONS[guessed[position]] == labels[i]else: right +=f"<|{labels[i]}|>"in written[position] said = normalise(re.sub(r"<\|[^|]*\|>", " ", written[position])) want = normalise(transcripts[i]) errors += edits(want, said) words +=len(want) emotion, wer = right /len(test_rows), errors / wordsdel model, head, optimiser, parameters gc.collect()if DEVICE.type=="mps": torch.mps.empty_cache()return emotion, werresults = {}print(f"{'design':>26}{'seed':>5}{'emotion':>9}{'word error':>11}{'minutes':>9}")for design in ("a classifier head", "a token in the vocabulary"): key ="head"if design.startswith("a classifier") else"token" scored = []for seed in SEEDS: started = time.monotonic() emotion, wer = run(key, seed) scored.append((emotion, wer))print(f"{design:>26}{seed:>5}{emotion:>9.3f}{wer:>11.3f} "f"{(time.monotonic() - started) /60:>9.1f}", flush=True) results[design] = scored
design seed emotion word error minutes
a classifier head 0 0.735 0.000 93.3
a classifier head 1 0.777 0.000 93.4
a classifier head 2 0.761 0.001 93.3
a token in the vocabulary 0 0.742 0.000 91.5
a token in the vocabulary 1 0.760 0.000 91.5
a token in the vocabulary 2 0.766 0.000 91.3
Ninety-odd minutes a run, and all six land between 0.735 and 0.777.
Show the code
print(f"{'design':>28}{'emotion':>22}{'word error rate':>24}")for design, runs in results.items(): emotion =sorted(r for r, _ in runs) wer =sorted(w for _, w in runs)print(f"{design:>28} "f"{np.mean(emotion):>8.3f} ({emotion[0]:.3f}-{emotion[-1]:.3f}) "f"{np.mean(wer):>10.3f} ({wer[0]:.3f}-{wer[-1]:.3f})")print(f"{'guessing':>28}{1/len(EMOTIONS):>8.3f}")
design emotion word error rate
a classifier head 0.758 (0.735-0.777) 0.000 (0.000-0.001)
a token in the vocabulary 0.756 (0.742-0.766) 0.000 (0.000-0.000)
guessing 0.167
Read the ranges in brackets before the means in front of them.
Conclusion
This experiment does not separate the two designs. The ranges overlap on both measures, and on a test set this size a few clips moves a mean further than the gap between them. Anyone reporting one seed here would have got a clean-looking result and it would have meant nothing.
What is not a measurement is where the work lands. A two-layer decoder in the token design has to produce the emotion and spell the word; in the head design it only spells, and the emotion comes off the encoder through its own layer. The emotion is in the encoder either way, so where you attach the readout does not change what there is to read — it changes how much the decoder is carrying.
Whatever that costs, it is a small-model concern. Whisper’s decoder is far larger and already emits several control tokens before it writes anything, so one more costs it nothing noticeable. The trap would have been to run this at one size, watch one design lose, and call the design worse when what I had measured was my own decoder being too small.
So the architectural argument stands on its own terms rather than on the numbers: a vocabulary entry instead of a reshaped head, one loss instead of two, and an emotion the transcription can actually see.