Emotional transitions — listen and judge

30 August 2026 · 399 clips from a study of how to make a voice turn from one feeling to another inside two sentences. Everything is explained below in plain language. The numbers say one thing; the point of this page is to find out whether your ears say the same.

Read this first: the measurements here are weak, and that is the main result. The scoring model we use to judge emotion can barely tell these clips apart. When the same sentence is generated once with the brief "reads as intense sadness" and once with "reads as intense amusement", the two differ by 0.12 of their own natural variation — and 0 of 30 texts reach even one full unit of separation. So every ranking below sits on an instrument that is nearly blind. Your ears are the better instrument, which is why this page exists.

1. What we were trying to do

A person who is grieving and then finds something bitterly funny does not switch. The grief is still under the amusement at the end of the sentence, and the amusement was already forming before the words arrived. Everything the model does today is a switch: brief the first sentence for sadness, brief the second for amusement, generate. There is a seam, and you can hear it.

The question was whether we can make the feeling move continuously — and, more importantly, move the way a body moves.

2. The three things we can push on

Everything below is built from three levers, all applied while the model speaks, without retraining anything.

leverwhat it ishow strong
Adapter (LoRA) a small patch of extra weights trained on the most extreme 1 % of examples for one feeling. You load it and turn a dial from 0 to about 1.5. moves emotion a little, delivery style a lot
Steering vector inside the model a sentence is a long list of numbers. Take the average of the angriest clips, subtract the average of middling ones, and you have a direction pointing at anger. Nudge the model along it while it speaks. the strongest of the three, but only in a narrow band
Guidance (CFG) run the model twice per 80 ms frame — once with the emotional instruction, once without — and exaggerate the difference. weakest, and costs 1.93× the time

The steering vector is where a fade naturally lives: it acts on the current moment only, so its weight can change from one 80 ms frame to the next. The adapter cannot be faded safely — changing its dial mid-sentence changes the weights that produced the memory the model is still reading from. So in everything below, adapters stay at a constant setting and the fade lives in the steering vectors and the prompt.

3. The conditions you are about to hear

Each clip is the same two sentences, the same voice, the same random seed. Only the method changes. Two of the fourteen are references and two are controls — they are there to tell you what "no transition" and "a meaningless transition" sound like, which is the only way to judge whether the rest are doing anything.

namewhat it doeswhy it is in the study
ANCHOR Athe whole clip briefed as the first feeling only reference: what "fully A" sounds like
ANCHOR Bthe whole clip briefed as the second feeling only reference: what "fully B" sounds like. If you cannot hear a difference between the two anchors, that is the headline result, not a mistake.
M0 — stepsentence one briefed A, sentence two briefed B, nothing fades what the system does today. The thing everything else has to beat
M1 — both onboth emotion adapters loaded at once, both feelings named the obvious first idea
M2 — linear crossfadetwo steering vectors, weights sliding linearly A→B the naive fade. Kept deliberately, because it has a flaw worth hearing
M3 — equal-power fadethe same fade, but the total push is held constant throughoutthe corrected version of M2
M4 — fade with a flooras M3, but A never fades below a quarter strength "the grief stays under the amusement"
M5 — anchored to the sentencethe turn is centred on the actual pause between the sentences, not on the halfway point of the clockreal turns happen at a moment, not uniformly
M6 — multi-ratethe body leads: arousal and tension turn ~0.3 s before the feeling does, and a "flatness" direction is subtracted throughout to keep the voice sounding realthe winner on the numbers. The hypothesis was that a transition is not one fade but several at different speeds
M7 — words onlyno vectors at all; the brief simply describes the arc the cheapest possible method, and the only one the model has seen in training
M8 — words + floor fadeM7 and M4 togetherdoes describing the arc help a fade?
M9 — three-way guidanceguidance with two competing instructions whose weights fadethe most expensive option, on a subset only
M10a further variant of the scheduleadditional shape
C RANDthe same amount of pushing, but on a shuffled schedule with no A→B storythe control that matters. If a method does not beat this, its "transition" is just perturbation

4. What the numbers found

1. The instrument is nearly blind. The two anchors differ by 0.12 of their own scatter; 0 of 30 texts separate. On whole clips only 3 of 40 emotion scores and 1 of 57 voice-quality scores differ at all, and the biggest single difference is on Affection — a feeling that appears in neither brief. This reproduces one of the oldest findings in the project: the emotional instruction in the prompt, on its own, does almost nothing.
2. Four methods beat the shuffled control. M2 (p = 0.001), M3 and M4 (p = 0.005) and M6 (p = 0.043) all move the measured feeling more than an identical amount of random pushing does. So the ordered A→B schedule is doing something real. M6 wins on the typical clip (2.53 against a "no transition" floor of 1.82).
3. But nothing beats the step. The best method against today's behaviour is M4, better on 20 of 30 texts, p = 0.099 — not significant. The margin over the control is about 0.8 where the measurement noise is 1.82. A real effect on a blunt instrument, not a solved problem.
4. The words-only method is the worst in the study. M7 scores below both the control and a clip held at a single feeling. This follows from finding 1: if the brief is not a lever, a method whose only lever is the brief has nothing to push with.

One prediction was right, two were wrong. Written down before the run: that the multi-rate method (M6) would win — right. That describing the arc in words would punch above its cost — wrong, and backwards. That keeping a floor of the old feeling would matter most — wrong; the method that actually ends with the old feeling still audible is M6, through asymmetry rather than a hard floor.

A technical note, in case you notice it: the naive crossfade (M2) does not push harder in the middle, as we first assumed — it pushes less. Two half-strength directions at a wide angle are shorter than one at full strength; for sadness against amusement, which sit almost opposite each other, the middle of M2 is at 50 % of the strength of its ends. M3 fixes that by construction. Whether you can hear the difference is one of the things we would like to know.

5. Listen

Pick a pair. Within each, the two anchors come first, then today's behaviour, then the methods, then the control. Rate anything you have an opinion about — 👍 if the turn sounds like one performance, 👎 if it sounds like two clips glued together or like nothing happens, and add a note if you want. Your ratings stay in this browser; the button at the bottom copies them all as text you can paste back.

6. What we would most like to know from you

  1. Can you hear the difference between ANCHOR A and ANCHOR B at all? The measurement says barely. If your ears disagree, the instrument is the problem and not the model — and that changes what we do next more than any other answer.
  2. Does M0 (the step) actually sound broken? It scores as "barely turns", but a seam can be obvious to a listener and invisible to a score.
  3. Does M6 sound like one performance where the others sound like two?
  4. Does M2 sound weaker in the middle than M3? That is a specific, measurable claim about the arithmetic, and the ear is the tiebreaker.
0 rated stored in this browser only

7. Your feedback, as text