Loop Forge

15 August 2026

MiniMax H3 with the Turbo LoRA

The Turbo LoRA setup for MiniMax H3: six steps instead of twenty plus, on a low VRAM card, entirely local in ComfyUI. The prompt carries a separate audio line, so the character says something and the model generates the speech along with the motion, from one still image.

Base model
minimax_h3_fl2va_pruned_int8_convrot
Text encoder
qwen3vl_32b_minimax_h3_nvfp4_awq
Turbo LoRA
minimax_h3_turbo_v4_step600_ema
Sampling
6 steps, simple scheduler
Watch on YouTube

MiniMax H3 + Turbo LoRA: image to video with sound on low VRAM. Opens on YouTube in a new tab.

What we found

Can you run MiniMax H3 on a 12GB GPU?

Yes. Everything local on this site ran on an RTX 3060 with 12GB of VRAM, including speech, singing and a one minute short film.

The Turbo LoRA brings H3 image to video down to six steps, and the spoken line still transcribes back word for word.

Reference to video at 864×480 ran out of memory past about 8 seconds on 12GB, and each generation took about 0.6 GPU-hours. A rented RTX 5090 ran the same prompts at 1344×768.

1 prompt

Reference images

Start image, The still the clip was generated from
Start imageThe still the clip was generated from

Materials and links