Text prompts steer genre, mood, and tempo. Nothing in the prompt layer steers
harmony — no key, no chords, no changes. Below is that gap, measured against
Stable Audio 3, and what we found trying to close it through init_audio.
Press play.
Three generations from Stable Audio 3. The text prompt and seed never change — the only variable is a THIRI-computed harmonic skeleton conditioning the render. Switch sources mid-playback and hear the difference at the same bar.
Clip B lands on the chart — including the G♭dim7 chromatic passing chord THIRI's reharmonization engine inserted at bar 8 (it shares three of four tones with F7; voice-leading score 0.86).
Clip C moves the harmony without touching the prompt — bars 9–12 swap to the tritone-substitution branch (C7 → G♭7). The vibe stays; the changes move.
WHAT THIS SHOWS — AND WHAT IT DOESN'T. The gap is real and measured:
the text-only control grades 2 of 12 bars against the chart, which is chance.
Prompting does not hold harmony. But B and C landing on their charts is not
evidence the model read one. A pre-registered attribution sweep found accuracy tracks
how much of the THIRI skeleton survives diffusion (r = +0.98); at init_noise
0.9, where the skeleton is gone, accuracy falls back to 2 of 12. Both skeletons also
grade 12 of 12 alone — so leakage alone reaches the observed ceiling. Audio conditioning
is the wrong wire for symbolic control.
Read the full negative result →
Stable Audio 3 renders in seconds, on licensed data, with editing pros can trust — the hard infrastructure problems, already solved. The next question arrives with every working musician who picks it up: “can it follow my changes?” We ran the measurement. Then we tried to close the gap through audio conditioning, and published the test that falsified that path.
Genre, mood, instrumentation, tempo — the prompt layer is genuinely good at feel. A chart is a different kind of instruction, and that's the piece we bring: prompts keep the vibe, THIRI holds the changes.
The major-label alliances put professional creators on the roadmap — and pros arrive with changes in hand. THIRI is how the model follows them, demonstrated live above.
The pro surface taking shape around Stable Audio speaks the DAW's language — chord tracks, MIDI, sessions. THIRI is that vocabulary as a service, already running as an API.
You solved speed. You solved trust. Control is the race that's left — and harmonic control is the hardest half of it.
Where THIRI fitsTHIRI is a chord-intelligence engine: analysis, voicings, voice-leading, and reharmonization computed from pitch-class-set theory — not sampled from a language model. It answers the same way every time, in milliseconds, from a production API that agents and plugins can call directly.
Any musician-facing Stable Audio surface needs a chord lane the moment a player touches it. THIRI is that lane as a service — suggestion, voicing, reharmonization — without building a theory engine in-house.
Our own attribution test says the honest version out loud: harmony passed as audio can't be separated from the audio it rode in on. The fix isn't a better skeleton — it's a pathway that accepts chord symbols directly. THIRI supplies the labeled harmonic ground truth to train exactly that.
Inpainting rebuilds the sound; THIRI context tells each regenerated bar what it must resolve to — so every edit stays harmonically true to the surrounding track.
Agents composing with Stable Audio need harmony they can trust. THIRI's MCP server plans the arc; the model renders it.
The attribution write-up lives on the training site. If you are building symbolic control for generative audio, the clips and the method are public.
dennison@bluesprincemedia.com · GitHub · MCP at mcp.thiri.ai