Well, I had seen many posts about Jev and there were many people that tried building a similar architecture for many different things. One of the coolest ones was PlayJev where they made it read a 448px game frame in a forward pass with no generated text, it was really cool
Honestly, since I was working in voice for a while now, initially I had thought it was a pretty natural use-case to try an audio-native Jev-like model for voice agents. Quick multiple decisions with useful probabilities? It’s a really neat tool to determine certain states in the conversation.
So, I tried building one. (a really small one, hope it works)
Well, I named it Prosodia. It can ingest speech/audio and can predict calibrated decisions without taking a look at the transcriptions. One clip gets encoded once by Whisper’s encoder and then every question branches over that, resulting with a probability distribution. I didn’t want to run the decoder to incorporate the transcriptions for the audio, because there are separate tools that’d let you extract such decisions from ASR+Audio combinations and I wanted this to be audio-native without implicitly relying on separate ASR results.
Try the live demo - record audio, edit questions and options at request-time, get probability distributions back.
TL;DR: Prosodia works as a direct speech decision pipeline, but my initial “prosody helps” result did not survive a stronger baseline. Whisper retained measurable pitch information, and the current pointer readout still fails one important Jev-like behavior (option interaction). The model is tiny (3.59M trainable params) and the dataset is small (~7 hours), so this is an exploration, not a benchmark.
What is Jev?
Jev isn’t a chat model. You hand it a state (some text, a situation, whatever context you have), a set of questions, and the allowed answers for each one, and it hands back a probability distribution per question. All of them at once, in a single forward pass, with no text generated anywhere.
The answers are typed: Choice (pick one of these options), Score (rate it on an ordered scale), Noul (is this true?). Instead of asking an LLM “is this ticket urgent, reply in JSON” and hoping, you just get p(urgent) = 0.42 back. Those numbers are supposed to be calibrated, when it says 0.42, it should be right about 42% of the time.
For voice agents, where you’re making quick judgements constantly and can’t afford a generation call for each one, this seemed like a natural fit.
Architecture
The shape is Jev’s: audio encoded once into a state, every question branches over that same state in structural isolation (questions are folded into the batch dimension, so cross-question attention is unrepresentable, not just masked), and a pointer readout scores each option against the branch vector. Typed outputs, Choice, Score, Noul, calibrated probabilities out, no autoregressive decode anywhere.
For encoders I tried four things:
- Whisper - frozen encoder, decoder never runs (the control that ended up winning everything)
- WavLM - frozen, self-supervised (the original primary)
- Explicit prosody - three scalars per frame at 50Hz: F0 via
librosa.pyin, RMS energy, voicing probability. Hand-built, deliberately the opposite of Jev’s strong-pretrained-representation premise, a diagnostic not a contender - Text-only - audio muted, dialogue history alone. Never sees the current utterance’s words
Everything trainable is a 3.59M-parameter stack learned from scratch, trained with cross-entropy + 0.5 x Brier.
I had some rules defined to actually produce some hardcoded GT labels, pitch was one of them and it was being captured by the Whisper’s encoder. So as a question pitch_direction, with options [rising, falling, level across the utterance] was able to be introduced to the dataset.
Pitch wasn’t alone though, there are four of these deterministic labels:
speaker: Choice over the six Friends leads, who cover ~83% of all utterancespitch_direction: falling / level / rising, from the F0 slope in semitones per second so it’s comparable across voicesspeaking_rate: slow / medium / fast, words-per-second binned into tertiles fit on train onlyloudness: quiet / moderate / loud, from RMS, same tertile treatment
Gold answers come from a fixed rule over the signal, so they’re unlimited, free, and self-verifying. Are they totally accurate? No, but I thought they provide enough signal to have some fun with the model after it’s trained. One caveat: speaking_rate requires a word count so it’s partially lexical, and for the prosody encoder these questions are circular (it gets F0/RMS as input, so it’d be reading the answer off its own features), they’re only a real test for Whisper and WavLM.
Where it differs from Jev
From what I understood, Jev’s options interact before choosing them. From Archer Hume’s independent probe of the Jev API, appending an irrelevant fifth option shifts mean log-odds between untouched options by about -0.28 across 10 randomized blocks, which proves Jev’s options are jointly encoded before the choice.
Quick precision on that: the -0.28 shift is the reverse-engineering probe’s result on real Jev, not mine. My model was never measured at -0.28; by the algebra below its shift is exactly zero. And the reason any nonzero shift matters: if each option were scored independently, an irrelevant newcomer mathematically could not move the relative odds between the other two. A shift means “the model considered the set”; zero means “it scored a list”.
The model readout computes:
logits = (W_branch @ h) @ (W_option @ o)^T
Each option’s logit is an independent inner product against its own frozen embedding. So p(a)/p(b) can’t change when you add another option, this satisfies independence of irrelevant alternatives exactly.
That’s exactly why the zero matters: this readout is provably in the functional class the Jev probe falsifies. My model has Jev’s interface but not its mechanism. The fix is options that attend to each other instead of independently-embedded frozen vectors. And no, Jev didn’t get option interaction from RLCD (that’s the unpublished calibration training recipe, a loss-side thing, while option interaction is architectural). NanoJev’s Choice head uses set attention over the candidates, which is the kind of mechanism that could produce interaction, though I haven’t verified whether theirs actually violates IIA either.
Is this thing actually something like Jev?
I thought a bit about this, I don’t know what RLCD actually is, and I had to remove option interactions as well, but I think given the data constraints this was one of the bests I could’ve come up with.
On the training part: by “same family as NanoJev” I mean neither of us has RLCD, nobody outside TypeSafe does, and both of us substituted strictly proper scoring rules on hard labels instead. Mine is cross-entropy + 0.5 x Brier; theirs is CE / Brier / paired-reward variants. Where we differ is the backbone (they finetune a 0.6B Qwen3 with differential learning rates; I freeze the audio encoder and train 3.6M params from scratch) and task weighting (they weight tasks explicitly; I didn’t, which later cost me a 13-point swing on an unrelated question when I removed 1.75% of labels, since unweighted joint training means every question tugs on one shared encoder).
So… where’s the data?
There is no “speech decisions” dataset sitting around, so I repurposed MELD: Friends utterances labeled with emotion and sentiment.
It’s ~13.7k clips total, ~10k for train and 2,610 for test. Clips are tiny (median ~2.5 seconds), so the whole corpus is roughly 7 hours of train, under 2 hours of test. That’s the entire world the model ever saw.
It fights you the whole way too: the splits are dialogue-disjoint but not speaker-disjoint (the same six Friends voices are in train and test, which flatters anything acoustic), ~2.5% of clips declare timestamps too short to contain their own transcripts, training caps audio at 30s because WavLM’s relative-position bias is O(T²), and my laptop OOM’d a few times before everything was working fine.
…so, that’s why the model wasn’t generalized well, I neither have the data nor the compute, as the gpu poor I am
Results
The model alone wasn’t so good and didn’t perform much better than the simple baselines I picked (i.e. just picking the majority labels), and didn’t perform better than the published text-only results on MELD.
Emotion classification
All encoders bracket the 0.481 majority accuracy. Macro-F1 is the only column showing any signal (whisper 0.221 vs majority 0.093). For scale, my best weighted-F1 is 0.437 from ~7 hours of audio and 3.6M trainable params from scratch, against published text-only MELD baselines around 0.57-0.60. Also note the architecture itself was never tested here, it’s constant across all twelve arms, so nothing in this grid bears on whether the Jev shape helps.
Expected Calibration Error
Well, ECE was well known and I heard it again in Jev (their headline calibration number is an ECE of 0.0313, so the choice was deliberate), so I reported it in here as well. It’s literally the average gap between how confident a model is and how often it’s right.
I added a simple baseline that only picks the majority of the options per question. It ignores the audio entirely and always emits the train class marginal, e.g. emotion becomes “neutral, 48% sure” on every single clip. Well, it didn’t work well since picking majority of classes nearly minimizes ECE on its own: calibration is a marginal property, it only asks whether stated confidence matches observed frequency on average, so a predictor that never distinguishes between inputs satisfies it almost perfectly.
So, my baseline got the most right (duh) and ECE was pretty much minimized. Train-fit constant at 0.0097 on both multi-class questions vs 0.0404 / 0.0213 for the best trained arm. I couldn’t use it to rank my experiments. That’s why I ranked my experiments by Brier instead. Brier is the mean squared error between your predicted probabilities and the one-hot truth, and it’s strictly proper: it’s minimized only by reporting what you actually believe, so a constant can’t game it. It just scores the marginal’s error, and anything that knows something about the input beats it.
Brier tells the actual story
Brier on emotion per arm, lower is better: whisper C-temp wins at 0.667, the train-fit constant sits at 0.710, and every prosody and text-only arm loses to the constant. Rank all twelve arms by ECE and by Brier and the orderings come out near-independent (Spearman +0.147 / -0.294 / -0.266 across the three questions). So, ECE wasn’t ranking calibration quality, it was ranking distance from a degenerate solution.
The… prosody effect?
The dataset I used (MELD) ships one label per utterance, I digged a bit and found EmotionLines (via the EmotionX 2019 release), the dataset MELD was built from. It had kept the raw annotator vote string, so I had actual annotator votes about the emotions. One correction though: the five annotators saw text only, so the vote spread (e.g. “2000030” over neutral/joy/sadness/fear/anger/surprise/disgust) measures transcript ambiguity, not prosody. MELD re-annotated with audio-visual and never released those votes, so an audio-side annotator distribution simply doesn’t exist here. Nothing more than the vote string, but it joined perfectly: 2,610 of 2,610 test rows on speaker+text.
Well, I tried running several encoders (the four from the Architecture section above), the one that ingests prosody explicitly didn’t perform well, hell, just knowing the dialogue history when the speech in the audio is somewhat clean was better. To be precise: Whisper did clearly beat history-only (emotion Brier 0.671 vs 0.712), so the lexical-content-in-latent-space story holds for Whisper; it’s the hand-built prosody arm that carried essentially nothing (0.718, worse than history alone).
I did get excited about a stratified version of this once, prosody looked better than the text baseline exactly where transcripts misled, then I scored every arm against its own stratum’s base rate instead (the clear stratum is 59.4% neutral, the ambiguous one 29.7%), and the effect died.
Skill vs each stratum’s own marginal, 0 = “as good as a constant”: prosody -0.110 / -0.025, never above zero anywhere; only whisper clears its base rate, and only where the transcript was already clear.
So, just the prosody, doesn’t carry enough info the transcription alone would destroy, at least in the dataset I’ve tried this out with.
I didn’t have enough data
Well, I believe the actual strength of such a model as Jev is essentially being able to learn a generalized-enough latent space, to handle different questions and options freely. I didn’t have enough data, some of the emotion classes sat near zero F1, why? well, there aren’t nearly enough data for some of them, that’s why.
Train counts: neutral 4710, joy 1743, surprise 1205, anger 1109, sadness 683, disgust 271, fear 268.
I tried some diagnostic runs with inverse-frequency class weights but yeah that also didn’t work well. The diagnostic actually split three ways: joy recovered for real (F1 0.216 to 0.310, recall 0.152 to 0.328, the samples were there all along, plain cross-entropy was suppressing the second-commonest class); sadness didn’t move at all (recall never above 2.9% on any encoder under a 6.9x upweight, at 683 examples that reads as the representation not holding the distinction, not the objective); and fear/disgust “recovered” at precision 0.018-0.051, i.e. 95-97% of those predictions wrong, recall bought at nothing. Plus collateral damage: the aggressive weighting broke surprise (0.314 to 0.132 on whisper).
The data… f#!@
I was just checking through some of the clips I had in the data, well, some of them were just too short to contain the transcript, so the audio/transcript alignment was off too.
The poster child is meld-test-173-2: “No, she doesn’t” in a declared span of 177 milliseconds. I can’t embed the actual Friends clip here (Warner Bros owns that audio; MELD’s GPL license can’t give it away), so here’s the same sentence recreated, first at natural length, then cut to the same 177ms the dataset declares:
full version:
cut version in the dataset (177ms):
So, I had a running whisper server on my local already, I just checked roundtrip ASR results to see if WER hits too high, a.k.a the ground truth transcript is just too different than what model predicts. WER blew up embarrassingly (worst case: ref “y’know” vs 33 ASR words, WER 33, on a one-word reference WER tells you nothing), and my replacement matcher anchored on the first word and reported 34.7% of clips at zero coverage while showing obviously-aligned pairs scoring 0.00. Third attempt with LCS ratio survived both failure modes: 40.7% of clips transcribe perfectly, median 0.89.
At least the corpus wasn’t actually too bad, only 3.3% of the samples had the alignment mismatch (plus 2% timestamp defects like the one above, now filtered from derived labels, and ~7% single interjections ASR just can’t score), so I just filtered them out.
final words
It was a cool thing to try out, and obviously it requires a way bigger dataset (preferably something longer than 7 hours :( ) to try it out.
I think it would be pretty nice to have a generalized version of this to act like a gateway in actual voice agents, though not sure how long we need STT->LLM->TTS loops considering full-duplex speech models getting tool-call support.
Code: alperiox/audio-jevlike · Demo: alperiox/prosodia · Trained on MELD (GPL-3.0).
