Loop Forge

5 September 2026

MiniMax H3 long takes from chained renders

MiniMax H3 renders top out at about fifteen seconds. To get past that, every render opens on an exact replay of the last 39 frames of the render before it, so the join is the same frames rather than similar frames. Character, street and voice all stay put across it.

References are handed in per render, so a new character or a new block of street can arrive halfway through a take. Twenty seven prompts across five takes, one per render, copied from what was sent.

Workflow
ComfyUI-H3-Continuous, chaining MiniMaxH3AddGuide
Handoff
39 frames plus audio from the end of each render
Tags
<Picture N> is positional: it comes from the order images are wired in, not from the prompt
Music
non_diegetic_music is N/A in all twenty seven; score goes on over the finished cut
Watch on YouTube

MiniMax H3 in ComfyUI: long videos from chained renders. Opens on YouTube in a new tab.

What we found

How do you make a MiniMax H3 video longer than 15 seconds?

H3 renders top out at about fifteen seconds. Chain them instead: each render opens on an exact replay of the last 39 frames of the one before, so the join is the same frames rather than similar ones.

Hand references in per render, and a new character or a new stretch of street can arrive partway through a take. The longest take here runs 61 seconds across seven renders.

Faces drift over six renders. A latent handover and a VAE handover were both tried, and neither moved the number.

27 prompts

Market Street walk, RTX 3060

One continuous front-tracking walk down Market Street in San Francisco, seven renders long.

Market Street walk, RTX 5090

The same seven shots on a rented RTX 5090, reworked for the longer shot that card could hold. It adds two instructions the local set does not carry: one holding the camera steady, one keeping her expression neutral between the scripted beats.

Mia talking to camera

Talking to camera, unbroken, across four renders.

Kate, six renders of speech

Six renders of continuous speech. This is the take where the face drifts. Each segment is sized to the words it has to carry, because runtime with nothing scripted in it is where the model invents speech-shaped filler.

Kate singing on a stage

Singing on a stage, three renders.

Reference images

Block 1, Aerial photograph of that block
Block 1Aerial photograph of that block
Amber, Character plate
AmberCharacter plate
Block 2, Aerial photograph of that block
Block 2Aerial photograph of that block
Sofia, Character plate
SofiaCharacter plate
Block 3, Aerial photograph of that block
Block 3Aerial photograph of that block
Block 4, Aerial photograph of that block
Block 4Aerial photograph of that block
Allie, Character plate
AllieCharacter plate
Block 5, Aerial photograph of that block
Block 5Aerial photograph of that block
Block 6, Aerial photograph of that block
Block 6Aerial photograph of that block
Dany, Character plate
DanyCharacter plate
Block 7, Aerial photograph of that block
Block 7Aerial photograph of that block
Mia, Three-panel character sheet
MiaThree-panel character sheet
Kate, Character plate
KateCharacter plate
Kate, The singer
KateThe singer
Stage, The stage
StageThe stage

Materials and links