Prosodia · live
Your clip is encoded once by Whisper's encoder, then every question below branches over that single encoding and is answered in parallel as a probability distribution. No transcript, no speech recognition, no generated text — Whisper's decoder never runs.
Edit freely — the model takes question text and option sets at request time, so these are not fixed. Options are comma-separated, minimum two. Each answer reports the nearest question the model was actually trained on: a low similarity means the answer is unsupported, not clever.
Computed from the signal — ◆ is the ground truth for your clip
Human-annotated categories — trained on acted sitcom affect, so expect drift on natural speech
Out of scope — the options are six sitcom characters, so it has to pick one even though your voice is none of them