Diffusion Single File
comfyui

very slow

#34
by Heouzen - opened

minimax_h3_ref2va_pruned_fp8_scaled.safetensors
qwen3vl_32b_minimax_h3_int8_convrot.safetensors
H200x1, 1920x1088, 2MP, 24 fps, 10 second, 20 steps.
[INFO] Requested to load MiniMaxH3
[INFO] loaded completely; 19984.52 MB loaded, full load: True
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 20/20 [46:45<00:00, 140.28s/it]
[INFO] Prompt executed in 00:49:26

The quality of this one kills LTX though

H200 🫣
with 1MP what is the generation time?

Have you used xformers? They are much faster that pytorch attention, just install them on top of pytorch (latest versions conflicts can be resolved by getting latest patched xformers if i remember correctly in Facebook github).
There's no other way to make it fast except reducing quality or making smaller picture size. For such huge resolution its very fast, i had same speed in full Kandinsky 5 Pro diffuser but for 640x368.

I don't see the reason to waste resources on audio generation if model not capable to generate a simple clip of live concert. I haven't tested H3 yet, but latest LTX is the "king of slop" in sloppiness, its like we returned in early Stable diffusion times.
LTX2.3 example (prompt: People shocked of Incredible concert of best classical musicians on Times Square live)

vlcsnap-2026-08-10-09h07m36s335

(prompt: clothes in the store shelf in the mall) LTX2.3

vlcsnap-2026-08-10-09h09m26s265

Sign up or log in to comment