TAU-HOME.COM
LOADING

Qwen3.8-Flash-Next NVFP4 Benchmarked on Single DGX Spark: 1M Context with FP8 KV Cache

A benchmark of Qwen3.8-Flash-Next NVFP4 on a single 128GB DGX Spark demonstrates 1M context with FP8 KV cache, 4-stream concurrency, and full multimodal input s

tau · October 7, 2026

#DGX Spark #Qwen3.8 #NVFP4 #FP8 KV #LocalLLM

Qwen3.8-Flash-Next NVFP4 Benchmarked on Single DGX Spark: 1M Context with FP8 KV Cache

On October 7, 2026, benchmark measurements and serving recipes were published for NVIDIA's official Qwen3.8-Flash-Next NVFP4 quantized checkpoint running on a single NVIDIA DGX Spark machine equipped with 128GB of unified memory. By combining 4-bit weights with an FP8 KV cache, the setup successfully demonstrates a 1-million-token (1M) context envelope alongside full multimodal text, image, and video ingestion on a single desktop workstation form factor.

Historically, operating frontier-class model architectures with million-token context windows required multi-GPU server clusters or rack-mounted enterprise hardware. The empirical data published for this Qwen3.8-Flash-Next NVFP4 runtime shows that carefully tuned unified-memory allocations and quantized KV caching can turn a standalone 128GB workstation into a capable, multi-tenant local AI server rather than a constrained hardware experiment.

FP8 KV Cache and 1M Context: Expanding Capacity to 1.43M Tokens

The most prominent technical milestone in this serving recipe is the dramatic expansion of usable context memory enabled by an FP8 KV cache.

Under conventional 16-bit (BF16) KV caching, maintaining long context windows quickly saturates the 128GB unified memory budget of a single DGX Spark. By adopting an FP8 KV cache format, the runtime substantially compresses memory footprint per sequence:

  • 1.43-Million-Token KV Pool: At target memory allocation, the configuration provides approximately 1,431,164 KV tokens, representing a 1.8x to 1.9x capacity expansion compared to baseline BF16 KV caching.
  • Enterprise-Scale Context Windows: With support for up to 1M tokens in active context, single-box deployments can ingest full codebases, comprehensive legal dossiers, multi-hour meeting transcripts, and extended browser-agent traces without context truncation.
  • Multi-Session Memory Headroom: The expanded cache capacity allows multiple independent agent sessions to retain persistent conversational state and shared tool history concurrently on the same physical host.

High-Throughput Prefill and 4-Stream Aggregate Scaling at 86 tok/s

In addition to sustained context depth, the runtime demonstrates balanced throughput characteristics across both prompt ingestion and generation:

  • Rapid Prefill Ingestion: In stress testing, the setup achieved approximately 1,495 tok/s when processing an extensive 400K-token prefill prompt under the FP8 KV cache profile. Shorter 32K-token prompts reached approximately 1,769 tok/s, maintaining a consistent 1,500 to 2,000 tok/s prefill ingestion envelope.
  • Single-Stream Generation: Single-stream prose generation achieved an interactive decoding rate of roughly 37 tok/s.
  • Four-Stream Concurrent Serving: When scaling to 4 concurrent streams, aggregate decoding throughput scaled to approximately 86 tok/s across the host.

This concurrency scaling illustrates an important operational pivot: on a dedicated local AI appliance, workloads frequently involve concurrent agent loops, automated code analysis tasks, and background document queries rather than a single user waiting on a single terminal prompt.

Multimodal Capabilities, Speculative Decoding Limits, and Operational Trade-offs

The runtime preserves full vision-language capabilities, accepting both static image inputs and dynamic video frames alongside textual prompts. However, the benchmark report notes several architectural constraints and practical trade-offs:

  • Speculative Decoding Degradation on Multimodal Inputs: Speculative decoding acceleration delivers diminished returns on multimodal requests because current draft models cannot directly process multimodal embeddings. Consequently, image and video queries experience reduced speculative drafting efficiency, falling back toward baseline non-speculative decoding speeds.
  • High Token Consumption of Video Frames: Video ingestion rapidly consumes token allocations during frame sampling. Even with a 1M-token ceiling, ingesting high-resolution or multi-minute video clips can exhaust the KV cache pool rapidly, requiring disciplined frame sampling intervals.
  • Workload-Dependent Recipe Selection: Alternative single-Spark recipes in the developer community prioritize peak single-stream decoding speeds over massive context depth. Infrastructure operators must balance their requirements between low-latency single-stream interactive generation and deep-context, multi-stream concurrency.

Sources