Loop Forge

23 September 2026

Qwen-Image 2.1, run locally

Qwen-Image 2.1 generates and edits with one 7 billion parameter model, and its autoencoder carries an alpha channel, so transparency comes out of the model rather than a background remover. Everything below ran locally on an RTX 3060, with every prompt exact and every timing measured.

Watch on YouTube

26 prompts

Transparent backgrounds

Editing by circling

One scene, four edits

Multiple references

Panoramic generation

Text rendering

Resolution and cost

Where it struggles

The model, and the machine

Transformer
7B single-stream DiT, 32 layers
Text encoder
Qwen3-VL 8B
VAE
64 channel RGBA, 16x spatial compression
Native output
2K, about 4.2 megapixels
Peak VRAM
10.5 to 11.7 GiB at int8
Licence
Qwen Research, non-commercial
Ran on
RTX 3060, 12GB (local)

Materials and links