TensorFold Extension Brings MiniMax H3 Video Generation to Apple Silicon Macs via MLX
An open-source extension integrates the MiniMax H3 video generation pipeline into TensorFold, an MLX-based Apple Silicon runtime. It delivers pre-built images a
On October 4, 2026, open-source developer keys (@u1tra_instinct) announced an extension that integrates MiniMax H3 video generation into TensorFold for Apple Silicon Macs, publishing a pull request to the upstream GitHub repository alongside a pre-built image for community testing.
![]()
Image source: https://github.com/ashhart/TensorFold
By extending TensorFold—an MLX-powered runtime originally built for memory-efficient MoE LLM inference—into joint multimodal video and audio diffusion, this release enables creators and engineers to experiment with local MiniMax H3 generation directly on Mac hardware rather than relying on remote cloud GPU clusters.
Bridging MLX Runtime TensorFold with MiniMax H3
TensorFold was originally developed by Ash Hart (@ashhart) as an exact-first local inference runtime on Apple's MLX framework. Its design focuses on serving large models under reduced resident memory pressure, allowing Apple Silicon systems to host architectures that typically overwhelm standard local setups.
This new extension adapts that core foundation to MiniMax H3's multimodal generation pipeline:
- Local Text-to-Video Execution: Enables prompt-driven generation of video with synchronized stereo audio directly on Apple Silicon unified memory without external server dependencies.
- Image-to-Video Expansion: Shortly after the initial release, the developer noted that image-conditioned video generation is actively being added to allow image-to-video workflows.
- Pre-Built Testing Image: A pre-built runtime image was provided alongside the pull request, allowing developers and early testers to run experiments without manual compilation steps.
Grounded in Open-Source Precedents: MLX and h3.c
MiniMax H3 is a 33B multimodal diffusion transformer paired with text encoders, demanding substantial memory bandwidth and computational throughput during multi-step denoising passes.
The developer explicitly framed the project on the shoulders of existing Apple Silicon research, crediting several foundational contributors across the ecosystem:
- antirez's h3.c and h3-metal: antirez's C and Metal-native MiniMax H3 inference engine and hardware-level Metal optimizations were explicitly credited as foundational prior work.
- Apple MLX Community Work: The broader MLX open-source ecosystem that pioneered accelerated machine learning execution on Apple Silicon served as a technical foundation.
- TensorFold Core Architecture: Ash Hart's original lightweight runtime served as the host framework, streamlining local serving workflows and reducing system overhead.
Release Maturity and Hardware Considerations
Before considering this tool for production workflows, developers should take note of several early-stage technical realities:
- Pull Request Stage: The code is currently submitted as a pull request against the upstream
ashhart/TensorFoldrepository, with iterative refinements, community feedback, and code cleanups ongoing. - Unified Memory Requirements: Due to the parameter footprint of MiniMax H3, local execution practically demands Apple Silicon hardware with high unified memory configurations, such as high-spec MacBook Pro or Mac Studio models.
- Ongoing Speed and Quality Benchmarking: The developer reported that a 5-way video comparison benchmark using the same prompt is in preparation alongside a video specialist, focusing on assessing visual output quality alongside inference speed.
Sources
- GitHub Repository: ashhart/TensorFold
- Developer Release Post: @u1tra_instinct on X (October 4, 2026)
- Reference Project: antirez/h3.c (h3-metal)