MiniMax H3 on One RTX 5090
On 3 September I wanted to answer three practical questions about MiniMax H3, also known as Hailuo 3.0. Could its open-weight video model run on one RTX 5090? How long would a clip take? Which of the available checkpoints and turbo LoRAs would be worth using?
The aim was not to rank several video models. These are configurations of the same model, changed through pruning, quantisation and step-reducing LoRAs. I wanted a usable configuration for this machine and measurements from the files I would actually run.
Test setup
The tests used ComfyUI’s official MiniMax H3 text-to-video workflow with stock settings. I did not enable SageAttention or any --fast options.
| Component | Configuration |
|---|---|
| GPU | RTX 5090, 32GB VRAM |
| CPU and RAM | Intel Core i9-13900K, 93GB usable RAM |
| Operating system | Ubuntu 24.04, NVIDIA driver 580.173, CUDA 13.0 |
| Runtime | ComfyUI 0.34.0, PyTorch 2.14.0 with CUDA 13.0 |
| Output | 1344 x 768, 124 frames, 24fps, 5.17 seconds, native stereo audio |
| Sampling | res_multistep, simple scheduler, normally 20 steps |
Every comparison used the same rooftop-chase prompt and seed 757358688076805. Wall time ran from submitting the workflow until the MP4 was saved locally, including model loading, generation and encoding. I sampled VRAM with nvidia-smi every two seconds.
What the settings mean
The names combine three separate choices: which version of the model is loaded, how its numbers are stored, and how many times it runs the generation loop.
| Term | Meaning here |
|---|---|
| Full and pruned | Two sizes of the H3 checkpoint. Pruned is the smaller model variant; this is separate from whether its weights use BF16, INT8 or another number format. |
| BF16 | A 16-bit floating-point format. It was the largest and least compressed option in these tests. |
| INT8 | 8-bit integer quantisation. Model values are represented with smaller integers and scale factors, reducing storage and memory at the cost of some approximation. |
| FP8 scaled | 8-bit floating-point values with scaling to make better use of their limited numeric range. |
| Mixed INT4/INT8 | A quantised checkpoint that combines 4-bit and 8-bit integer storage instead of using one precision throughout. |
| NVFP4 | NVIDIA’s 4-bit floating-point format for Blackwell GPUs. The 4 is the number of bits used for each E2M1 value; small blocks also carry scale information so useful range is retained. |
| LoRA | A small low-rank adapter applied to a base model. The turbo LoRAs here are designed to make usable output with only four or eight generation steps. |
| Steps | The number of denoising iterations, not the number of video frames. At each step the model refines the latent video towards the prompt. More steps mean more passes through the expensive part of the model, so they usually take longer. |
This matters when reading the timings. The 4-step run does one fifth as many denoising iterations as a 20-step run, but it uses a turbo LoRA built for that shorter process. It is not simply the standard model stopped early. Loading and encoding also take time, so wall time does not fall in exact proportion to the step count.
Outputs
All four versions followed the requested structure: a rooftop sprint, a jump, the landing and another launch. The dusk city, pursuers, fog and flying traffic also appeared consistently. The four-step version was brighter and more saturated, with larger and less controlled flying cars, but it remained coherent enough for prompt iteration.
Measurements
| Configuration | Resolution | Steps | Wall time | Peak VRAM |
|---|---|---|---|---|
| Pruned INT8, cold | 864 x 480 | 20 | 105.2s | 31.1GB |
| Pruned INT8 | 1344 x 768 | 20 | 260.4s | 31.3GB |
| Pruned FP8 scaled | 1344 x 768 | 20 | 326.8s | 28.0GB |
| Pruned INT8 plus turbo LoRA | 1344 x 768 | 8 | 160.3s | 28.7GB |
| Pruned INT8 plus 768p turbo LoRA | 1344 x 768 | 4 | 70.1s | 31.3GB |
| Full INT8 | 1344 x 768 | 20 | 340.5s | 28.5GB |
| Pruned NVFP4 | 1344 x 768 | 20 | 260.4s | 27.0GB |
| Pruned mixed INT4/INT8 | 1344 x 768 | 20 | 281.2s | 27.9GB |
| Full BF16 plus INT8 text encoder | 1344 x 768 | 20 | OOM at 86GB RSS | RAM limit |
| Full BF16 plus NVFP4 text encoder | 1344 x 768 | 20 | OOM at 78GB RSS | RAM limit |
| Pruned BF16, cold | 1344 x 768 | 20 | 505.9s | 28.2GB |
Pruned INT8 produced 5.17 seconds of 768p video in 260.4 seconds. NVFP4 finished in the same time while reducing peak VRAM from 31.3GB to 27GB. Its checkpoint was also 12.5GB rather than 21GB, which made it the practical choice for this card.
More precision did not buy an obvious improvement in this example. Full INT8 took 340.5 seconds, while pruned BF16 took 505.9 seconds. The full BF16 checkpoint failed twice. ComfyUI staged the 63GB state dictionary in system memory, and the process was killed after reaching 86GB and 78GB of resident memory. The limit was the machine’s 93GB of RAM, not the GPU’s 32GB of VRAM.
Resolution affected time almost perfectly in proportion to pixel count. Moving from 864 x 480 to 1344 x 768 increased the number of pixels by 2.49 times and wall time by 2.48 times. Peak VRAM barely changed because the model itself dominated memory use.
What I would use
For draft renders, I would use the four-step turbo LoRA. It produced a preview in 70.1 seconds, which is fast enough to test composition and prompt changes without waiting 4.3 minutes each time. Once the prompt worked, I would render the selected version with pruned NVFP4 at 20 steps. The final render is still roughly 50 times slower than realtime, but it runs locally on one consumer GPU with native audio.
This was one prompt and one seed, with visual differences judged by eye. The clips show what each configuration produced for this storyboard; they do not support a general quality ranking. The useful result is narrower: H3 runs on the RTX 5090, NVFP4 is the sensible 20-step checkpoint for this setup, and the four-step LoRA makes iteration much less tedious.
Next
These were stock-setting numbers. The next useful test would keep the prompt, seed and NVFP4 checkpoint fixed, then measure ComfyUI’s --fast options and SageAttention separately. That would show which optimisations save time without making the output visibly worse.