Project Maya: Up to 16 GPU Scaling and RTX 3090 / V100 Speed Boosts
Project Maya rolls out multi-GPU scaling up to 16 GPUs, boosting RTX 3090 decode speeds and Tesla V100 prefill performance up to 674 tok/s.
A new release of Project Maya, an open-source high-speed inference framework for large language models (LLMs), has been officially rolled out in October 2026. This update focuses on multi-GPU distributed scaling and significant decode and prefill speed optimizations tailored to widely deployed workstation and enterprise GPUs.
Project Maya is an inference acceleration engine designed to run large language models efficiently on local and on-premises hardware. Developed with community contributions (acknowledging collaborators such as @needmorevram), this release expands cluster scaling capabilities up to 16 GPUs while delivering targeted performance enhancements across popular NVIDIA architectures.
Scaling Up to 16 GPUs for Distributed Inference
The primary architectural enhancement in this release is expanded parallel GPU support, allowing setups scaling up to 16 GPUs.
- 16 GPU Scale-Out: Supports larger GPU cluster topologies, enabling practitioners to shard and host high-parameter models across multi-GPU or multi-node infrastructures.
- Distributed Compute Efficiency: Optimizes communication and tensor parallel execution pipelines to maintain high utilization during large-model inference workloads.
RTX 3090 Decode Acceleration and Tesla V100 Prefill Optimizations (Up to 674 tok/s)
This release delivers measurable benchmark gains across both consumer flagship and enterprise data center GPUs:
- NVIDIA GeForce RTX 3090 Decode Boost: Accelerates token generation (decode) throughput on 24GB VRAM RTX 3090 setups.
- NVIDIA Tesla V100 Prefill Performance:
- 1x Tesla V100: Prompt prefill throughput reaches up to 620 tok/s.
- 2x Tesla V100: Scaled prefill performance reaches up to 674 tok/s, significantly cutting down prompt evaluation latency.
- 24GB VRAM Large Model (GLM) Optimizations: Tailored memory and computation tuning on 24GB cards specifically aimed at larger-footprint models such as GLM, improving overall decode and prefill throughput.
Hardware Scope and Implementation Notes
When adopting Project Maya's latest release, operators should note the following hardware considerations:
- Heterogeneous GPU Topologies: Support for mixed-architecture or heterogeneous GPU clustering remains unconfirmed in the current release.
- Non-NVIDIA Hardware Limits: Support for non-NVIDIA GPUs (such as AMD Radeon RX 6750 XT) remains limited, with primary optimization benefits currently centered on NVIDIA CUDA environments.
Sources
- GitHub Repository: mw00/project-maya
- Peasant Smith on X: Project Maya New Release Announcement