How do you get speech and singing from a video model?
MiniMax H3 and LTX 2.5 both generate audio in the same pass as the picture. Write the spoken line into the prompt and the model produces the speech along with the lip movement, from one still image.
Turn the model's own music off in the prompt and score in the edit. Separately generated music will not line up across cuts or joins.
Size each render to the words it has to carry. Runtime with nothing scripted in it is where the model invents speech-shaped filler, and on speech the two models fail differently: one hallucinates dialogue nobody asked for, the other returns silence.



















