Local AI Runtime Benchmarks: Comparing exllamav3, llama.cpp, vLLM, and TensorFold Across 16GB to 256GB Memory Classes
Empirical benchmark results comparing local LLM inference speeds (tok/s) and Time to First Token (TTFT) across RTX 5060 Ti (16GB), RTX 5090 (32GB), and GB10 (12
AI engineer and researcher Niklas Lenz (@niklaslenz_ai) has released the benchmark findings from his “Local AI Test Part 2”, measuring empirical token generation throughput (tok/s) and Time to First Token (TTFT) across four hardware memory classes (16GB to 256GB) using leading local inference engines—exllamav3, llama.cpp, vLLM, and TensorFold.

Image source: @niklaslenz_ai on X
Expanding upon an earlier 5-setup evaluation, Part 2 tests 9 distinct serving configurations across 4 memory classes using an identical set of 10 application-building prompt tasks. The new benchmark integrates optimized kits and recipes contributed by @MiaAI_lab and @ashxhart, granting up to 150 minutes of execution time for the three heaviest application workloads.
Benchmark Architecture and Evaluation Metrics
The test matrix spans consumer desktop GPUs through multi-box unified memory hardware.
Hardware and Model Matrix
- 16 GB (RTX 5060 Ti): Qwen3.8-27B — Mia's exllamav3 kit vs. Niklas's llama.cpp setup
- 32 GB (RTX 5090): Qwen3.8-27B — Mia's vLLM recipe vs. Ash's TensorFold configuration
- 128 GB (Single GB10 Box): Qwen3.8-Flash-Next — Mia's TensorFold recipe vs. Mia's vLLM kit
- 256 GB (2× GB10 Boxes): GLM-5.3-Flash — Mia's vLLM kit vs. Mia's TensorFold kit vs. Ash's TensorFold recipe
Measured Metrics
- Throughput (tok/s): Median output tokens per second per request.
- Latency (TTFT): p50 Time to First Token across all requests.
- Cold vs. Warm State:
Coldindicates disabled prefix caching or the initial request in a multi-turn task;Warmdenotes an active prefix-cache hit on subsequent requests.
1. 16GB Tier (RTX 5060 Ti) — Qwen3.8-27B
On a 16GB VRAM budget, running the 27B parameter Qwen3.8 requires aggressive quantization.
- Mia's kit (exllamav3 1.4.4, EXL3 2.5 bpw)
- Output Speed: 56.34 tok/s (median)
- TTFT: 3.98 s (p50)
- Niklas's setup (llama.cpp b11151, UD-Q3_K_XL ~3.8 bpw)
- Output Speed: 39.71 tok/s (median)
- TTFT: 2.51 s (p50)
exllamav3 (2.5 bpw) achieved 56.34 tok/s in raw decoding throughput, outpacing llama.cpp by approximately 42%. Conversely, llama.cpp (UD-Q3_K_XL) demonstrated lower initial prompt processing latency with a TTFT of 2.51 seconds.
2. 32GB Tier (RTX 5090) — Qwen3.8-27B
On a flagship 32GB GPU, the benchmark compared 4-bit serving pipelines under vLLM and TensorFold.
- Mia's recipe (vLLM 0.27.1, NVFP4)
- Output Speed: 75.43 tok/s (median)
- TTFT: 20.39 s (prefix cache off, all requests cold)
- @ashxhart's recipe (TensorFold 0.6.0, MLX 4bit)
- Output Speed: 352.28 tok/s (median)
- TTFT: Cold 30.01 s / Warm 0.84 s
TensorFold 0.6.0 with MLX 4bit delivered a substantial 352.28 tok/s decode rate. While un-cached cold starts incurred high initial latency (20.39s on vLLM and 30.01s on TensorFold), TensorFold dropped to a responsive 0.84-second TTFT once prefix cache hits took effect.
3. 128GB Tier (Single GB10 Box) — Qwen3.8-Flash-Next
In a 128GB unified memory environment, Mia's two serving configurations were evaluated on the Qwen3.8-Flash-Next architecture.
- TensorFold 0.3.6.3 (MLX 4bit)
- Output Speed: 105.5 tok/s (median)
- TTFT: Task initial request (Cold) 4.99 s / Subsequent requests (Warm) 0.82 s
- vLLM kit (NVFP4)
- Output Speed: 72.8 tok/s (median)
- TTFT: Task initial request (Cold) 2.12 s / Subsequent requests (Warm) 1.21 s
TensorFold maintained a 45% throughput lead in sustained generation (105.5 tok/s vs. 72.8 tok/s). While vLLM handled the initial prompt faster on the first request (2.12s vs. 4.99s), TensorFold achieved lower latency (0.82s) across follow-up iterations.
4. 256GB Tier (2× GB10 Boxes) — GLM-5.3-Flash
Across a dual-box 256GB deployment, three setups were tested with the larger GLM-5.3-Flash model.
- Mia's vLLM kit (vLLM, EXL3 4bpw)
- Output Speed: 54.43 tok/s (median)
- TTFT: 2.62 s (Warm)
- Mia's TensorFold kit (TensorFold 0.6.0, EXL3 4bpw)
- Output Speed: 89.45 tok/s (median)
- TTFT: 1.13 s (Warm)
- @ashxhart's recipe (TensorFold 0.6.0, MLX 4bit)
- Output Speed: 55.66 tok/s (median)
- TTFT: 1.19 s (Warm)
When comparing the identical EXL3 4bpw quantization format, TensorFold 0.6.0 (89.45 tok/s, TTFT 1.13s) outperformed vLLM (54.43 tok/s, TTFT 2.62s) by 64% in throughput while offering lower warm latency. Within TensorFold, the EXL3 4bpw recipe also yielded higher throughput than the MLX 4bit variant (55.66 tok/s).
Multi-Tier Benchmark Summary Table
| Memory Class | Hardware Setup | Model | Runtime & Quantization | Throughput (tok/s) | TTFT Latency | Recipe Source |
|---|---|---|---|---|---|---|
| 16 GB | RTX 5060 Ti | Qwen3.8-27B | exllamav3 1.4.4 (EXL3 2.5 bpw) | 56.34 | 3.98 s | Mia's kit |
| 16 GB | RTX 5060 Ti | Qwen3.8-27B | llama.cpp b11151 (UD-Q3_K_XL ~3.8 bpw) | 39.71 | 2.51 s | Niklas setup |
| 32 GB | RTX 5090 | Qwen3.8-27B | vLLM 0.27.1 (NVFP4) | 75.43 | 20.39 s (Cold) | Mia's recipe |
| 32 GB | RTX 5090 | Qwen3.8-27B | TensorFold 0.6.0 (MLX 4bit) | 352.28 | Cold 30.01 s / Warm 0.84 s | @ashxhart |
| 128 GB | 1× GB10 | Qwen3.8-Flash-Next | TensorFold 0.3.6.3 (MLX 4bit) | 105.50 | Cold 4.99 s / Warm 0.82 s | Mia's recipe |
| 128 GB | 1× GB10 | Qwen3.8-Flash-Next | vLLM kit (NVFP4) | 72.80 | Cold 2.12 s / Warm 1.21 s | Mia's kit |
| 256 GB | 2× GB10 | GLM-5.3-Flash | vLLM (EXL3 4bpw) | 54.43 | 2.62 s (Warm) | Mia's kit |
| 256 GB | 2× GB10 | GLM-5.3-Flash | TensorFold 0.6.0 (EXL3 4bpw) | 89.45 | 1.13 s (Warm) | Mia's kit |
| 256 GB | 2× GB10 | GLM-5.3-Flash | TensorFold 0.6.0 (MLX 4bit) | 55.66 | 1.19 s (Warm) | @ashxhart |
Original source
- Niklas Lenz on X (@niklaslenz_ai): Local AI Test Part 2 Benchmark Results