Speech Emotion Recognition with Whisper: Classifier Head vs Vocabulary Token

audio
speech
architecture

Two designs over CREMA-D, three seeds each, and what changes when the label is a class index instead of a word the decoder has to say.

Author

Rosh Beed

Published

July 9, 2026

Whisper transcribes. It does not tell you how something was said, and how is often the point: the same sentence read angrily and read sadly is one transcript and two different messages.

The obvious route is a classifier.

Three lines, and it works.

I built it the other way, because of something Whisper already does.

Whisper’s Token Prefix

Whisper's multitask training format: a token sequence branching through previous-text prompting, a language tag, transcribe versus translate, and timestamp options.

A Whisper transcript does not start with words. It starts with a structured token prefix, and that prefix is what lets one model do transcription, translation and language identification without being three models.

Those are not a separate mechanism bolted on. They are ordinary vocabulary entries, predicted by the same softmax and trained by the same cross-entropy as any word. Look at the ids above:

<|startoftranscript|><|en|><|transcribe|><|notimestamps|>Hello, my name is Bes.<|endoftext|>

[50258, 50259, 50359, 50363, 15947, 11, 452, 1315, 307, 8190, 13, 50257]

The first four are the language and the task. Nothing distinguishes them from 15947 except what the model learned to do when it sees them.

So: is emotion a different kind of thing from language and task, or the same kind of thing? It is hard to argue it is different. Adding <emotion_happy> and its siblings to that vocabulary lets the model say what it heard through machinery it already has.

Show the code
import re

import numpy as np
import torch
import torch.nn as nn
from huggingface_hub import hf_hub_download
from transformers import WhisperForConditionalGeneration, WhisperProcessor

DEVICE = torch.device("mps" if torch.backends.mps.is_available() else "cpu")

MODEL = "openai/whisper-small"
BATCH, LR, EPOCHS, SEEDS = 4, 1e-5, 3, (0, 1, 2)

CREMAD = "3a5d5fa8a9945e122258fb6c52c243126186c227"
data = np.load(hf_hub_download("roshbeed/ai-residency-blog-data", "speech/cremad.npz",
                               repo_type="dataset", revision=CREMAD),
               allow_pickle=True)
offsets, labels = data["offsets"], list(data["labels"])
splits, EMOTIONS = list(data["splits"]), [str(e) for e in data["emotions"]]
# Read the audio out once. np.load on a compressed .npz is lazy, so indexing
# data["audio"] inside the loop decompresses the whole array on every iteration.
# Slices of the one array, not copies of it. The conversion to float happens a
# batch at a time, where it is needed.
samples = data["audio"]
clips = [samples[offsets[i]:offsets[i + 1]] for i in range(len(labels))]

# CREMA-D is twelve fixed sentences said by ninety-one actors. The sentence is in
# the filename, which is how the transcript is recovered.
SENTENCES = {
    "IEO": "It's eleven o'clock", "TIE": "That is exactly what happened",
    "IOM": "I'm on my way to the meeting", "IWW": "I wonder what this is about",
    "TAI": "The airplane is almost full", "MTI": "Maybe tomorrow it will be cold",
    "IWL": "I would like a new alarm clock", "ITH": "I think I have a doctor's appointment",
    "DFA": "Don't forget a jacket", "ITS": "I think I've seen this before",
    "TSI": "The surface is slick", "WSI": "We'll stop in a couple of minutes",
}
FILES = "180ad0a0636f15b42fb349d77a00758a3343d847"
names = np.load(hf_hub_download("roshbeed/ai-residency-blog-data",
                                "speech/cremad-files.npz",
                                repo_type="dataset", revision=FILES),
                allow_pickle=True)["files"]
transcripts = [SENTENCES[str(f).rsplit("/", 1)[-1].split("_")[1]] for f in names]

print(f"{len(clips):,} clips from CREMA-D, {len(EMOTIONS)} emotions: "
      f"{', '.join(EMOTIONS)}")
print(f"{splits.count('train'):,} train, {splits.count('test'):,} test")
print(f"guessing the emotion: {1 / len(EMOTIONS):.3f}")
7,442 clips from CREMA-D, 6 emotions: anger, disgust, fear, happy, neutral, sad
5,209 train, 1,117 test
guessing the emotion: 0.167

Two designs, one model. Both start from the same released whisper-small weights; the only difference is where the emotion comes out.

The head design hangs a classifier off the encoder. The token design adds six entries to the vocabulary and makes the decoder emit one before it writes the transcript, so the emotion is predicted by the same softmax as every other token and trained by the same loss.

Show the code
processor = WhisperProcessor.from_pretrained(MODEL)
EMOTION_TOKENS = [f"<|{e}|>" for e in EMOTIONS]


def build(design, seed):
    """One model, either with a classifier head or with six new vocabulary rows."""
    torch.manual_seed(seed)
    tokeniser = WhisperProcessor.from_pretrained(MODEL).tokenizer
    # `generate` forces <|en|><|transcribe|><|notimestamps|> after the start token.
    # The tokenizer emits neither the language nor the task by default, so a target
    # built without them trains the decoder against a prefix that generation never
    # produces. Transcription survives it; a single token at a fixed position does
    # not, because the position it was trained at is one generation never reaches.
    tokeniser.set_prefix_tokens(language="en", task="transcribe")
    model = WhisperForConditionalGeneration.from_pretrained(MODEL)
    head = None
    if design == "token":
        tokeniser.add_tokens(EMOTION_TOKENS, special_tokens=True)
        model.resize_token_embeddings(len(tokeniser))
    else:
        head = nn.Linear(model.config.d_model, len(EMOTIONS))
    return model.to(DEVICE), (head.to(DEVICE) if head else None), tokeniser


SOT = processor.tokenizer.convert_tokens_to_ids("<|startoftranscript|>")


def targets(tokeniser, indices, design):
    """Whisper prepends its own start token, so the tokenizer's is dropped."""
    if design == "token":
        text = [f"<|{labels[i]}|>{transcripts[i]}" for i in indices]
    else:
        text = [transcripts[i] for i in indices]
    ids = tokeniser(text, return_tensors="pt", padding=True).input_ids
    assert ids[0, 0].item() == SOT
    return ids[:, 1:].to(DEVICE)


def features_for(indices):
    audio = [clips[i].astype(np.float32) / 32768.0 for i in indices]
    return processor(audio, sampling_rate=16000,
                     return_tensors="pt").input_features.to(DEVICE)


train_rows = [i for i, s in enumerate(splits) if s == "train"]
test_rows = [i for i, s in enumerate(splits) if s == "test"]
print(f"{len(train_rows):,} training clips, {len(test_rows):,} held out")
print(f"the token design adds {len(EMOTION_TOKENS)} rows to a "
      f"{len(processor.tokenizer):,}-entry vocabulary")
5,209 training clips, 1,117 held out
the token design adds 6 rows to a 51,865-entry vocabulary

Training differs only in where the emotion sits. The token design puts it in the sequence and uses one cross-entropy over everything. The head design takes it out of the sequence and adds a second loss on the classifier.

Two numbers come back for each: how often the emotion is right, and how much of the transcript survives. The whole question is what the second costs.

Show the code
import gc
import time

def normalise(text):
    return re.sub(r"[^a-z' ]", "", text.lower()).split()


def edits(reference, hypothesis):
    previous = list(range(len(hypothesis) + 1))
    for i, want in enumerate(reference, 1):
        current = [i]
        for j, got in enumerate(hypothesis, 1):
            current.append(min(previous[j] + 1, current[j - 1] + 1,
                               previous[j - 1] + (want != got)))
        previous = current
    return previous[-1]


def run(design, seed):
    model, head, tokeniser = build(design, seed)
    parameters = list(model.parameters()) + (list(head.parameters()) if head else [])
    optimiser = torch.optim.AdamW(parameters, lr=LR)
    generator = torch.Generator().manual_seed(seed)

    for _ in range(EPOCHS):
        order = torch.randperm(len(train_rows), generator=generator).tolist()
        for start in range(0, len(order) - BATCH + 1, BATCH):
            batch = [train_rows[i] for i in order[start:start + BATCH]]
            features = features_for(batch)
            out = model(input_features=features, labels=targets(tokeniser, batch, design),
                        output_hidden_states=(head is not None))
            loss = out.loss
            if head is not None:
                pooled = out.encoder_last_hidden_state.mean(1)
                wanted = torch.tensor([EMOTIONS.index(labels[i]) for i in batch],
                                      device=DEVICE)
                loss = loss + nn.functional.cross_entropy(head(pooled), wanted)
            optimiser.zero_grad()
            loss.backward()
            optimiser.step()

    model.eval()
    right = errors = words = 0
    for start in range(0, len(test_rows), 16):
        batch = test_rows[start:start + 16]
        features = features_for(batch)
        with torch.no_grad():
            tokens = model.generate(features, max_new_tokens=48,
                                    language="en", task="transcribe")
            if head is not None:
                pooled = model.model.encoder(features).last_hidden_state.mean(1)
                guessed = head(pooled).argmax(-1).tolist()
        written = tokeniser.batch_decode(tokens, skip_special_tokens=False)
        for position, i in enumerate(batch):
            if head is not None:
                right += EMOTIONS[guessed[position]] == labels[i]
            else:
                right += f"<|{labels[i]}|>" in written[position]
            said = normalise(re.sub(r"<\|[^|]*\|>", " ", written[position]))
            want = normalise(transcripts[i])
            errors += edits(want, said)
            words += len(want)
    emotion, wer = right / len(test_rows), errors / words
    del model, head, optimiser, parameters
    gc.collect()
    if DEVICE.type == "mps":
        torch.mps.empty_cache()
    return emotion, wer


results = {}
print(f"{'design':>26} {'seed':>5} {'emotion':>9} {'word error':>11} {'minutes':>9}")
for design in ("a classifier head", "a token in the vocabulary"):
    key = "head" if design.startswith("a classifier") else "token"
    scored = []
    for seed in SEEDS:
        started = time.monotonic()
        emotion, wer = run(key, seed)
        scored.append((emotion, wer))
        print(f"{design:>26} {seed:>5} {emotion:>9.3f} {wer:>11.3f} "
              f"{(time.monotonic() - started) / 60:>9.1f}", flush=True)
    results[design] = scored
                    design  seed   emotion  word error   minutes
         a classifier head     0     0.735       0.000      93.3
         a classifier head     1     0.777       0.000      93.4
         a classifier head     2     0.761       0.001      93.3
 a token in the vocabulary     0     0.742       0.000      91.5
 a token in the vocabulary     1     0.760       0.000      91.5
 a token in the vocabulary     2     0.766       0.000      91.3

Ninety-odd minutes a run, and all six land between 0.735 and 0.777.

Show the code
print(f"{'design':>28} {'emotion':>22} {'word error rate':>24}")
for design, runs in results.items():
    emotion = sorted(r for r, _ in runs)
    wer = sorted(w for _, w in runs)
    print(f"{design:>28} "
          f"{np.mean(emotion):>8.3f} ({emotion[0]:.3f}-{emotion[-1]:.3f}) "
          f"{np.mean(wer):>10.3f} ({wer[0]:.3f}-{wer[-1]:.3f})")
print(f"{'guessing':>28} {1 / len(EMOTIONS):>8.3f}")
                      design                emotion          word error rate
           a classifier head    0.758 (0.735-0.777)      0.000 (0.000-0.001)
   a token in the vocabulary    0.756 (0.742-0.766)      0.000 (0.000-0.000)
                    guessing    0.167

Read the ranges in brackets before the means in front of them.

Conclusion

This experiment does not separate the two designs. The ranges overlap on both measures, and on a test set this size a few clips moves a mean further than the gap between them. Anyone reporting one seed here would have got a clean-looking result and it would have meant nothing.

What is not a measurement is where the work lands. A two-layer decoder in the token design has to produce the emotion and spell the word; in the head design it only spells, and the emotion comes off the encoder through its own layer. The emotion is in the encoder either way, so where you attach the readout does not change what there is to read — it changes how much the decoder is carrying.

Whatever that costs, it is a small-model concern. Whisper’s decoder is far larger and already emits several control tokens before it writes anything, so one more costs it nothing noticeable. The trap would have been to run this at one size, watch one design lose, and call the design worse when what I had measured was my own decoder being too small.

So the architectural argument stands on its own terms rather than on the numbers: a vocabulary entry instead of a reshaped head, one loss instead of two, and an emotion the transcription can actually see.

The full project is on GitHub.