Listens to a backing track and generates a matching, steerable solo as real audio.
Generative audioTransformersPyTorchEnCodec
What it is
SoloMuse listens to a backing track and generates a solo that fits it, as actual audio. Unlike a black-box model, it plans the performance first (when to play, how dense, how tense, how loud), so the output can be steered.
How it works
Situation layer. Reads the backing track (tempo, loudness, harmony/chroma, spectral features) into a compact 32-number summary.
Intent planner. A causal Transformer plans the musical intent 10 times per second: play or rest, note density, register, tension, dynamics, onsets and phrasing. You can override it, e.g. force more tension.
Renderer. An autoregressive multi-codebook Transformer language model generates Meta EnCodec audio tokens (4 codebooks × 1,024 tokens at 75 Hz), which are decoded into a waveform.