23 September 2026
Qwen-Image 2.1, run locally
Qwen-Image 2.1 generates and edits with one 7 billion parameter model, and its autoencoder carries an alpha channel, so transparency comes out of the model rather than a background remover. Everything below ran locally on an RTX 3060, with every prompt exact and every timing measured.
26 prompts
Transparent backgrounds
Editing by circling
One scene, four edits
Multiple references
Panoramic generation
Text rendering
Resolution and cost
Where it struggles
The model, and the machine
- Transformer
- 7B single-stream DiT, 32 layers
- Text encoder
- Qwen3-VL 8B
- VAE
- 64 channel RGBA, 16x spatial compression
- Native output
- 2K, about 4.2 megapixels
- Peak VRAM
- 10.5 to 11.7 GiB at int8
- Licence
- Qwen Research, non-commercial
- Ran on
- RTX 3060, 12GB (local)

























