Writing

MiniMax H3 on One RTX 5090

· AI · Local AI, Video Generation

On 3 September I wanted to answer three practical questions about MiniMax H3, also known as Hailuo 3.0. Could its open-weight video model run on one RTX 5090? How long would a clip take? Which of the available checkpoints and turbo LoRAs would be worth using?

The aim was not to rank several video models. These are configurations of the same model, changed through pruning, quantisation and step-reducing LoRAs. I wanted a usable configuration for this machine and measurements from the files I would actually run.

Test setup

The tests used ComfyUI’s official MiniMax H3 text-to-video workflow with stock settings. I did not enable SageAttention or any --fast options.

ComponentConfiguration
GPURTX 5090, 32GB VRAM
CPU and RAMIntel Core i9-13900K, 93GB usable RAM
Operating systemUbuntu 24.04, NVIDIA driver 580.173, CUDA 13.0
RuntimeComfyUI 0.34.0, PyTorch 2.14.0 with CUDA 13.0
Output1344 x 768, 124 frames, 24fps, 5.17 seconds, native stereo audio
Samplingres_multistep, simple scheduler, normally 20 steps

Every comparison used the same rooftop-chase prompt and seed 757358688076805. Wall time ran from submitting the workflow until the MP4 was saved locally, including model loading, generation and encoding. I sampled VRAM with nvidia-smi every two seconds.

What the settings mean

The names combine three separate choices: which version of the model is loaded, how its numbers are stored, and how many times it runs the generation loop.

TermMeaning here
Full and prunedTwo sizes of the H3 checkpoint. Pruned is the smaller model variant; this is separate from whether its weights use BF16, INT8 or another number format.
BF16A 16-bit floating-point format. It was the largest and least compressed option in these tests.
INT88-bit integer quantisation. Model values are represented with smaller integers and scale factors, reducing storage and memory at the cost of some approximation.
FP8 scaled8-bit floating-point values with scaling to make better use of their limited numeric range.
Mixed INT4/INT8A quantised checkpoint that combines 4-bit and 8-bit integer storage instead of using one precision throughout.
NVFP4NVIDIA’s 4-bit floating-point format for Blackwell GPUs. The 4 is the number of bits used for each E2M1 value; small blocks also carry scale information so useful range is retained.
LoRAA small low-rank adapter applied to a base model. The turbo LoRAs here are designed to make usable output with only four or eight generation steps.
StepsThe number of denoising iterations, not the number of video frames. At each step the model refines the latent video towards the prompt. More steps mean more passes through the expensive part of the model, so they usually take longer.

This matters when reading the timings. The 4-step run does one fifth as many denoising iterations as a 20-step run, but it uses a turbo LoRA built for that shorter process. It is not simply the standard model stopped early. Loading and encoding also take time, so wall time does not fall in exact proportion to the step count.

Outputs

Pruned INT8, 20 steps: 260.4 seconds, 31.3GB peak VRAM.
Pruned NVFP4, 20 steps: 260.4 seconds, 27.0GB peak VRAM.
Pruned INT8 with four-step turbo LoRA: 70.1 seconds, 31.3GB peak VRAM.
Pruned BF16, 20 steps: 505.9 seconds, 28.2GB peak VRAM.

All four versions followed the requested structure: a rooftop sprint, a jump, the landing and another launch. The dusk city, pursuers, fog and flying traffic also appeared consistently. The four-step version was brighter and more saturated, with larger and less controlled flying cars, but it remained coherent enough for prompt iteration.

Measurements

ConfigurationResolutionStepsWall timePeak VRAM
Pruned INT8, cold864 x 48020105.2s31.1GB
Pruned INT81344 x 76820260.4s31.3GB
Pruned FP8 scaled1344 x 76820326.8s28.0GB
Pruned INT8 plus turbo LoRA1344 x 7688160.3s28.7GB
Pruned INT8 plus 768p turbo LoRA1344 x 768470.1s31.3GB
Full INT81344 x 76820340.5s28.5GB
Pruned NVFP41344 x 76820260.4s27.0GB
Pruned mixed INT4/INT81344 x 76820281.2s27.9GB
Full BF16 plus INT8 text encoder1344 x 76820OOM at 86GB RSSRAM limit
Full BF16 plus NVFP4 text encoder1344 x 76820OOM at 78GB RSSRAM limit
Pruned BF16, cold1344 x 76820505.9s28.2GB

Pruned INT8 produced 5.17 seconds of 768p video in 260.4 seconds. NVFP4 finished in the same time while reducing peak VRAM from 31.3GB to 27GB. Its checkpoint was also 12.5GB rather than 21GB, which made it the practical choice for this card.

More precision did not buy an obvious improvement in this example. Full INT8 took 340.5 seconds, while pruned BF16 took 505.9 seconds. The full BF16 checkpoint failed twice. ComfyUI staged the 63GB state dictionary in system memory, and the process was killed after reaching 86GB and 78GB of resident memory. The limit was the machine’s 93GB of RAM, not the GPU’s 32GB of VRAM.

Resolution affected time almost perfectly in proportion to pixel count. Moving from 864 x 480 to 1344 x 768 increased the number of pixels by 2.49 times and wall time by 2.48 times. Peak VRAM barely changed because the model itself dominated memory use.

What I would use

For draft renders, I would use the four-step turbo LoRA. It produced a preview in 70.1 seconds, which is fast enough to test composition and prompt changes without waiting 4.3 minutes each time. Once the prompt worked, I would render the selected version with pruned NVFP4 at 20 steps. The final render is still roughly 50 times slower than realtime, but it runs locally on one consumer GPU with native audio.

This was one prompt and one seed, with visual differences judged by eye. The clips show what each configuration produced for this storyboard; they do not support a general quality ranking. The useful result is narrower: H3 runs on the RTX 5090, NVFP4 is the sensible 20-step checkpoint for this setup, and the four-step LoRA makes iteration much less tedious.

Next

These were stock-setting numbers. The next useful test would keep the prompt, seed and NVFP4 checkpoint fixed, then measure ComfyUI’s --fast options and SageAttention separately. That would show which optimisations save time without making the output visibly worse.

← All posts