Local AI Runtime Benchmarks: Comparing exllamav3, llama.cpp, vLLM, and TensorFold Across 16GB to 256GB Memory Classes

Empirical benchmark results comparing local LLM inference speeds (tok/s) and Time to First Token (TTFT) across RTX 5060 Ti (16GB), RTX 5090 (32GB), and GB10 (12

tau · October 4, 2026

#LocalLLM #vLLM #llamacpp #exllamav3 #TensorFold #Qwen38 #GPUBenchmark

Local AI Runtime Benchmarks: Comparing exllamav3, llama.cpp, vLLM, and TensorFold Across 16GB to 256GB Memory Classes

AI engineer and researcher Niklas Lenz (@niklaslenz_ai) has released the benchmark findings from his “Local AI Test Part 2”, measuring empirical token generation throughput (tok/s) and Time to First Token (TTFT) across four hardware memory classes (16GB to 256GB) using leading local inference engines—exllamav3, llama.cpp, vLLM, and TensorFold.

Benchmark infographic comparing local AI runtimes including exllamav3, llama.cpp, vLLM, and TensorFold across 16GB to 256GB VRAM hardware tiers

Image source: @niklaslenz_ai on X

Expanding upon an earlier 5-setup evaluation, Part 2 tests 9 distinct serving configurations across 4 memory classes using an identical set of 10 application-building prompt tasks. The new benchmark integrates optimized kits and recipes contributed by @MiaAI_lab and @ashxhart, granting up to 150 minutes of execution time for the three heaviest application workloads.

Benchmark Architecture and Evaluation Metrics

The test matrix spans consumer desktop GPUs through multi-box unified memory hardware.

Hardware and Model Matrix

  • 16 GB (RTX 5060 Ti): Qwen3.8-27B — Mia's exllamav3 kit vs. Niklas's llama.cpp setup
  • 32 GB (RTX 5090): Qwen3.8-27B — Mia's vLLM recipe vs. Ash's TensorFold configuration
  • 128 GB (Single GB10 Box): Qwen3.8-Flash-Next — Mia's TensorFold recipe vs. Mia's vLLM kit
  • 256 GB (2× GB10 Boxes): GLM-5.3-Flash — Mia's vLLM kit vs. Mia's TensorFold kit vs. Ash's TensorFold recipe

Measured Metrics

  • Throughput (tok/s): Median output tokens per second per request.
  • Latency (TTFT): p50 Time to First Token across all requests.
  • Cold vs. Warm State: Cold indicates disabled prefix caching or the initial request in a multi-turn task; Warm denotes an active prefix-cache hit on subsequent requests.

1. 16GB Tier (RTX 5060 Ti) — Qwen3.8-27B

On a 16GB VRAM budget, running the 27B parameter Qwen3.8 requires aggressive quantization.

  • Mia's kit (exllamav3 1.4.4, EXL3 2.5 bpw)
    • Output Speed: 56.34 tok/s (median)
    • TTFT: 3.98 s (p50)
  • Niklas's setup (llama.cpp b11151, UD-Q3_K_XL ~3.8 bpw)
    • Output Speed: 39.71 tok/s (median)
    • TTFT: 2.51 s (p50)

exllamav3 (2.5 bpw) achieved 56.34 tok/s in raw decoding throughput, outpacing llama.cpp by approximately 42%. Conversely, llama.cpp (UD-Q3_K_XL) demonstrated lower initial prompt processing latency with a TTFT of 2.51 seconds.

2. 32GB Tier (RTX 5090) — Qwen3.8-27B

On a flagship 32GB GPU, the benchmark compared 4-bit serving pipelines under vLLM and TensorFold.

  • Mia's recipe (vLLM 0.27.1, NVFP4)
    • Output Speed: 75.43 tok/s (median)
    • TTFT: 20.39 s (prefix cache off, all requests cold)
  • @ashxhart's recipe (TensorFold 0.6.0, MLX 4bit)
    • Output Speed: 352.28 tok/s (median)
    • TTFT: Cold 30.01 s / Warm 0.84 s

TensorFold 0.6.0 with MLX 4bit delivered a substantial 352.28 tok/s decode rate. While un-cached cold starts incurred high initial latency (20.39s on vLLM and 30.01s on TensorFold), TensorFold dropped to a responsive 0.84-second TTFT once prefix cache hits took effect.

3. 128GB Tier (Single GB10 Box) — Qwen3.8-Flash-Next

In a 128GB unified memory environment, Mia's two serving configurations were evaluated on the Qwen3.8-Flash-Next architecture.

  • TensorFold 0.3.6.3 (MLX 4bit)
    • Output Speed: 105.5 tok/s (median)
    • TTFT: Task initial request (Cold) 4.99 s / Subsequent requests (Warm) 0.82 s
  • vLLM kit (NVFP4)
    • Output Speed: 72.8 tok/s (median)
    • TTFT: Task initial request (Cold) 2.12 s / Subsequent requests (Warm) 1.21 s

TensorFold maintained a 45% throughput lead in sustained generation (105.5 tok/s vs. 72.8 tok/s). While vLLM handled the initial prompt faster on the first request (2.12s vs. 4.99s), TensorFold achieved lower latency (0.82s) across follow-up iterations.

4. 256GB Tier (2× GB10 Boxes) — GLM-5.3-Flash

Across a dual-box 256GB deployment, three setups were tested with the larger GLM-5.3-Flash model.

  • Mia's vLLM kit (vLLM, EXL3 4bpw)
    • Output Speed: 54.43 tok/s (median)
    • TTFT: 2.62 s (Warm)
  • Mia's TensorFold kit (TensorFold 0.6.0, EXL3 4bpw)
    • Output Speed: 89.45 tok/s (median)
    • TTFT: 1.13 s (Warm)
  • @ashxhart's recipe (TensorFold 0.6.0, MLX 4bit)
    • Output Speed: 55.66 tok/s (median)
    • TTFT: 1.19 s (Warm)

When comparing the identical EXL3 4bpw quantization format, TensorFold 0.6.0 (89.45 tok/s, TTFT 1.13s) outperformed vLLM (54.43 tok/s, TTFT 2.62s) by 64% in throughput while offering lower warm latency. Within TensorFold, the EXL3 4bpw recipe also yielded higher throughput than the MLX 4bit variant (55.66 tok/s).

Multi-Tier Benchmark Summary Table

Memory ClassHardware SetupModelRuntime & QuantizationThroughput (tok/s)TTFT LatencyRecipe Source
16 GBRTX 5060 TiQwen3.8-27Bexllamav3 1.4.4 (EXL3 2.5 bpw)56.343.98 sMia's kit
16 GBRTX 5060 TiQwen3.8-27Bllama.cpp b11151 (UD-Q3_K_XL ~3.8 bpw)39.712.51 sNiklas setup
32 GBRTX 5090Qwen3.8-27BvLLM 0.27.1 (NVFP4)75.4320.39 s (Cold)Mia's recipe
32 GBRTX 5090Qwen3.8-27BTensorFold 0.6.0 (MLX 4bit)352.28Cold 30.01 s / Warm 0.84 s@ashxhart
128 GB1× GB10Qwen3.8-Flash-NextTensorFold 0.3.6.3 (MLX 4bit)105.50Cold 4.99 s / Warm 0.82 sMia's recipe
128 GB1× GB10Qwen3.8-Flash-NextvLLM kit (NVFP4)72.80Cold 2.12 s / Warm 1.21 sMia's kit
256 GB2× GB10GLM-5.3-FlashvLLM (EXL3 4bpw)54.432.62 s (Warm)Mia's kit
256 GB2× GB10GLM-5.3-FlashTensorFold 0.6.0 (EXL3 4bpw)89.451.13 s (Warm)Mia's kit
256 GB2× GB10GLM-5.3-FlashTensorFold 0.6.0 (MLX 4bit)55.661.19 s (Warm)@ashxhart

Original source