
Reference-audio style conditioning
Handing the model ten seconds of the texture you want removes the ambiguity that adjectives like 'warm' or 'vintage' always carry.
MiniMax
Style conditioning from a reference clip rather than a text description.
In short
Figures verified 2026-08-10. This field moves quickly — re-check before relying on them.
Reference-audio conditioning is the reason this generation is still worth knowing about. Describing a sound in words is lossy — 'warm', 'vintage' and 'lo-fi' mean different things to different models, and to different people — while handing over ten seconds of the texture you want removes the ambiguity entirely. If you have a reference track whose production you are chasing, this is a more direct route than iterating on adjectives. The obvious caution is that conditioning on a commercial recording you do not own creates a derivative-work question that is unsettled and jurisdiction-dependent; conditioning on your own material does not. Full-song length and low cost make it usable for volume work, and while 2.5 supersedes it on raw quality, 1.5 remains the version people reach for when style matching matters more than vocal polish.
Strengths

Handing the model ten seconds of the texture you want removes the ambiguity that adjectives like 'warm' or 'vintage' always carry.

Full-length output rather than a clip means the reference style is applied across an entire arrangement, not just an intro.

Low per-generation cost makes it practical to try the same reference against several different lyric sheets.
How it compares
MiniMax Music 1.5 generates full-length songs and can be conditioned on a reference audio clip, steering style from an example rather than from adjectives. It generates at low cost per track.
Compare with MusicGenerate
Searched as minimax music 1.5 and minimax reference audio. Style conditioning from a clip is the reason to pick this version over a newer one.
Handing the model a short audio clip to steer style, instead of describing the style in words. Adjectives like 'warm' or 'vintage' mean different things to different models; ten seconds of the texture you want removes the ambiguity.
Use 1.5 when style matching from a reference matters more than vocal polish. Use 2.5 for better raw quality on a text-and-lyrics workflow.
Technically yes, legally no. Conditioning on a commercial recording you do not own raises a derivative-work question that is unsettled and jurisdiction-dependent. Use your own material when the output is going to be released.
Yes, full-length rather than clips, so the reference style is applied across an entire arrangement instead of just an intro.
Per-generation cost sits well below the Western premium tier, which is what makes trying the same reference against several lyric sheets practical.
A short excerpt is enough to carry style — ten seconds of the texture you want conveys more than a paragraph of adjectives.
Keep exploring
More models
Models and platforms we track, compared on capability and licence.
Style conditioning from a reference clip rather than a text description.