TAU-HOME.COM
LOADING

NVIDIA Open-Sources 'PixelUMM': An Encoder-Free Unified Multimodal Model in Raw Pixel Space

NVIDIA researchers (nv-tlabs) have open-sourced the code and model weights for PixelUMM, an encoder-free multimodal architecture that eliminates both VAEs and V

tau · October 5, 2026

#NVIDIA #PixelUMM #OpenSource #Multimodal #ComputerVision #nv-tlabs

NVIDIA Open-Sources 'PixelUMM': An Encoder-Free Unified Multimodal Model in Raw Pixel Space

On October 5, 2026, NVIDIA researchers (nv-tlabs) led by Cong Wei (@CongWei1230) officially open-sourced the code, model weights, and research paper for 'PixelUMM'—a unified multimodal foundation model that performs both visual understanding and generation directly in raw pixel space without relying on pretrained vision encoders or latent compression models.

Conventional multimodal vision models (VLMs and LMMs) almost universally decouple visual processing: they rely on pretrained Vision Transformer (ViT) encoders to extract high-level semantic features for comprehension, and Variational Autoencoders (VAEs) to compress pixels into latent representations for generative diffusion. PixelUMM eliminates these intermediate proxy components entirely, establishing an encoder-free paradigm where a single unified architecture consumes and generates raw visual pixels end-to-end.

Discarding ViTs and VAEs: A Unified Encoder-Free Architecture in Pixel Space

In existing vision-language systems, visual understanding and visual generation have traditionally been treated as distinct engineering tasks. Understanding workloads necessitated frozen ViT backbones that patch and summarize visual tokens, while generation pipelines depended on discrete or continuous latent spaces produced by autoencoders.

  • Unified Single-Network Pipeline: PixelUMM tokenizes raw image and video pixels directly, bypassing both external vision encoders and latent compression autoencoders.
  • Joint Understanding and Generation: Instead of chaining disjoint specialized models, the single architecture natively handles both visual reasoning tasks and image or video synthesis.
  • Preventing Encoder Bottlenecks and Information Loss: By avoiding the lossy compression of VAEs and the inductive semantic filtering of pretrained ViTs, the model retains fine-grained, low-level visual information directly from raw pixel signals.

This encoder-free formulation fundamentally streamlines the preprocessing and feature-extraction stack in multimodal systems, offering a unified representation of visual data that simplifies training and downstream model inspection.

Video Tokenization Ablation: Spatial vs. Temporal Efficiency and the p32/t1 Configuration

Operating directly in pixel space introduces significant scaling challenges for video data due to the rapid growth of spatio-temporal tokens. To address this, the researchers conducted detailed video tokenization ablation experiments evaluating how spatial and temporal dimensions should be balanced under an identical token budget of 1,024 pixels per token.

  • Experimental Setup: Evaluated token configurations allocating 1,024 pixels into each video token, contrasting two distinct design strategies: p16/t4 versus p32/t1.
  • p16/t4 Configuration: Combined a $16 \times 16$ spatial patch across 4 temporal frames ($16 \times 16 \times 4 = 1,024$ pixels).
  • p32/t1 Configuration: Combined a $32 \times 32$ spatial patch across a single temporal frame ($32 \times 32 \times 1 = 1,024$ pixels).
  • Ablation Findings: The p32/t1 configuration achieved the lowest training loss in the ablation benchmark, outperforming the temporally compressed p16/t4 configuration.

These findings suggest that when training multimodal models directly in raw pixel space, prioritizing high-fidelity spatial representation per token is more effective at minimizing reconstruction and training loss than aggressive temporal compression, providing an empirical foundation for encoder-free video tokenization.

Open-Source Impact, Auditability, and Production Considerations

By releasing the official code repository on GitHub alongside checkpoints on Hugging Face and the complete paper, the research team enables the machine learning community to inspect and audit the architecture directly.

  • Direct Inspection of the Inference Stack: With the public availability of the code repository (nv-tlabs/PixelUMM) and model weights (nvidia/PixelUMM), developers can verify and reproduce the end-to-end inference pipeline without managing separate ViT weights or autoencoder dependencies.
  • Computational and Memory Footprint: Because raw pixel inputs are processed without prior latent compression, processing high-resolution video streams or long contextual sequences demands careful memory management and substantial compute resources.
  • Serving and Deployment Requirements: As a research-grade open-source release, production deployments will require dedicated single-GPU and multi-GPU runtime optimizations, memory-efficient attention kernels, and specialized serving stacks tailored to large-scale pixel-space inference.

Sources