MiniMax H3 2-Pass Latent Upscale Workflow v2.0 for ComfyUI: 7s Video in 60s on RTX 5090

The Refmod + Reference 2-Pass Latent Upscale Workflow v2.0 for MiniMax H3 has been released on Civitai, generating 1MP 7-second videos in approximately 60 secon

tau · October 5, 2026

#MiniMaxH3 #ComfyUI #LatentUpscale #AIVideo #RTX5090

MiniMax H3 2-Pass Latent Upscale Workflow v2.0 for ComfyUI: 7s Video in 60s on RTX 5090

A new high-speed rendering setup for MiniMax H3, titled 'Minimax H3 - Refmod + Reference - 2 Pass Latent Upscale Workflow' v2.0, has been published on Civitai by creator R@aiaicreate (@aiaicreate), providing local creators with a rapid, reference-guided video generation pipeline within ComfyUI.

ComfyUI node graph interface showing the MiniMax H3 2-pass latent upscale workflow v2.0 for reference-to-video generation

Image source: X @aiaicreate / Civitai

Designed for local workstation deployment, the updated workflow demonstrates a measured production milestone on NVIDIA's flagship GeForce RTX 5090, rendering a 1-megapixel (1MP, approximately 1376×768) 7-second video clip in roughly 60 seconds while preserving visual fidelity and prompt alignment.

Mechanics of the 2-Pass Latent Upscale and v2.0 Enhancements

Standard generative video upscaling workflows typically rely on decoding low-resolution frames back into pixel space via a VAE, followed by standard image super-resolution and a subsequent VAE re-encoding pass. This round-trip process introduces substantial computational latency, consumes massive memory bandwidth, and frequently causes temporal flickering or drift away from the original reference subject across frames.

The 2-pass latent upscale workflow engineered by R@aiaicreate resolves these structural limitations by keeping the enlargement pass inside latent tensor space rather than incurring repetitive VAE conversions.

  • Pass 1 (Base Generation): Synthesizes initial video frames at lower native resolution using MiniMax H3's Reference-to-Video (R2V) capabilities and reference model (Refmod), locking in motion dynamics, subject appearance, and scene composition.
  • Pass 2 (Latent Upscale Pass): Feeds the uncompressed latent video representation directly into a dedicated latent upscaler, expanding the tensor to 1MP resolution without intermediate pixel decode cycles.
  • v2.0 Version Highlights: According to the release documentation on Civitai, the author designates v2.0 as the "fastest possible workflow" while substantially improving visual consistency with the initial reference image compared to earlier iterations.

RTX 5090 Benchmark Results and Rendering Efficiency

Foundation video models with temporal attention layers often demand minutes or even tens of minutes to render brief clips on consumer hardware. Achieving consistent 1MP or 1080p outputs locally has historically posed severe compute hurdles for daily production.

In verified testing on a desktop GeForce RTX 5090, the workflow recorded notable efficiency metrics:

  • Target Hardware: NVIDIA GeForce RTX 5090 GPU
  • Output Resolution and Duration: 1MP (1376×768 widescreen) / 7-second video clip
  • Total Rendering Duration: Approximately 60 seconds
  • Fidelity Profile: Reference-guided identity preservation without pixel-level drift or seam artifacts across sequential frames

By bypassing full VAE round-trips and leveraging the parallel tensor throughput of the RTX 5090, the generation cycle drops to near one minute. This drastically shortens the iteration feedback loop for creators exploring prompt variations, camera trajectories, and reference styling in production settings.

ComfyUI Node Setup and Practical Implementation Considerations

Creators downloading the workflow package from Civitai should ensure prerequisite dependencies and model placements are verified before launching the generation graph in ComfyUI:

  1. Custom Node Dependencies: Complex video graphs require dedicated node extensions for MiniMax H3 and latent upscaling pipelines. Community feedback on social threads noted missing custom node references such as NvidiaDLSSImageSequenceUpscale; users should execute an automated missing node check through ComfyUI-Manager prior to queuing the workflow.
  2. Latent Upscaler Model Weights: The latent upscaling stage relies on dedicated weights (such as h3_clean_latent_upscaler_film_epoch200.safetensors). These model checkpoints must be placed inside the appropriate models directory (e.g., ComfyUI/models/h3_latent_upscaler or configured path) to avoid runtime initialization failures.
  3. VRAM and Resolution Scaling: While optimized for the memory capacity and compute cores of the RTX 5090, creators running 24GB GPUs (such as the RTX 4090) should adjust base latent resolutions and upscale sampling steps to stay safely within physical VRAM boundaries.

Original source