Prosodia · live

An audio Jev

Your clip is encoded once by Whisper's encoder, then every question below branches over that single encoding and is answered in parallel as a probability distribution. No transcript, no speech recognition, no generated text — Whisper's decoder never runs.

This model was trained on acted sitcom audio. Your voice through a laptop mic is out of domain, so treat the affect answers sceptically. Speaker is meaningless for you — the option set is six Friends characters, so it must pick one. The acoustic questions carry a ◆ marking the true answer, computed from your clip by the same fixed rule used in training.

Questions

Edit freely — the model takes question text and option sets at request time, so these are not fixed. Options are comma-separated, minimum two. Each answer reports the nearest question the model was actually trained on: a low similarity means the answer is unsupported, not clever.

idle