TAU-HOME.COM
LOADING

Swift 1.5 GSQ-RCO GGUF Release Packs 27B-Class Qwen3.8 for 16GB GPUs

UkisAI's reasoning-efficient Swift 1.5 Qwen3.8-27B gets an ultra-light 2-3-bit GSQ-RCO GGUF packaging for llama.cpp, plus IQ3_S 16GB benchmark reports.

tau · October 8, 2026

#Qwen3.8 #UkisAI-Swift-1.5 #GGUF #llama.cpp #Quantization

Swift 1.5 GSQ-RCO GGUF Release Packs 27B-Class Qwen3.8 for 16GB GPUs

UkisAI has released ultra-light GSQ-RCO GGUF quantized builds of Swift 1.5 Qwen3.8-27B, its 27B-class reasoning-efficient model, ready to run in llama.cpp environments. The GSQ-RCO-GGUF release has been tracked as a public model on OpenModelStats since September 24, 2026, and advertises a compact 2-3-bit mixed-precision configuration.

A 27B-class open model running locally on a modest 16GB desktop GPU setup, evoking compact GGUF quantization for llama.cpp

Image source: ukisai / Hugging Face

The key point is that the model's reasoning efficiency and this quantization packaging are two separate stages. Swift 1.5 is UkisAI's derivative of Qwen3.8-27B that claims roughly 58.5% fewer thinking tokens and a score about 0.35% higher than the base in its Hugging Face model card, while GSQ-RCO-GGUF is the llama.cpp packaging of that Swift 1.5 model.

How Swift 1.5 and the GSQ-RCO-GGUF Release Relate

Swift 1.5 Qwen3.8-27B is described by UkisAI as a reasoning-efficient derivative of Qwen3.8-27B, claiming about 58.5% fewer thinking tokens with a score roughly 0.35% higher than the base. Those are UkisAI's marketing claims from the model card, not independently verified figures.

The GSQ-RCO-GGUF release is the quantization package built on top of Swift 1.5. According to the Hugging Face release page, it is a compact mixed-precision GGUF refined for Swift using the per-tensor allocations from ISTA-DASLab's GSQ-RCO release, targets llama.cpp, and includes the MTP head. Model-level benchmarks and quantization measurements are distinct, so the Swift 1.5 score claims should not be read as the quality of any specific quantization tier.

  • Base model: Swift 1.5 Qwen3.8-27B (UkisAI's reasoning-efficient derivative of Qwen3.8-27B)
  • This release: Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF (compact 2-3-bit mixed-precision, for llama.cpp)
  • Public tracking since: September 24, 2026 (per OpenModelStats)

IQ3_S Measured Figures and Confirmed Limits

Third-party aggregation (llm-bench.io, as of October 2026) states that the IQ3_S-mtp combination recorded up to about 51.2 tok/s of local inference speed, the best of 5 community runs. Within that page's aggregation structure, this is the peak for a specific hardware, tool, and quantization combination, not a single representative speed for the model.

A separate hands-on benchmark note (NemoClaw) states that it evaluated the IQ3_S quantization's performance, memory, reasoning, and coding tasks on an Ubuntu setup with an RTX 2000 Ada (16GB VRAM) GPU. The verifiable scope of that note is that the evaluation was performed; this article does not quote individual task scores as confirmed figures.

  • Limited to verified figures: the thinking-token and score claims in the HF model card, llm-bench.io's IQ3_S-mtp peak of about 51.2 tok/s (best of 5 runs), and the fact that NemoClaw performed its 16GB-environment evaluation
  • Excluded figures: the original X post's body was only partially retrievable (cut off at "9.18x spee…"), so the decode-speed figure the post claimed could not be independently verified and is not cited here

Who This Matters For

For users aiming to run a 27B-class model locally on a 16GB VRAM GPU, this release matters because it provides a llama.cpp-compatible IQ3_S path. The HF release table lists the IQ3_S file size at roughly 11.77GB, which confirms it sits in a size range that can be attempted on a 16GB setup. Actual perceived speed, context, and quality still depend on the user's GPU, llama.cpp runtime version, and quantization tier choice, so comparing the per-combination table on llm-bench.io against your own environment before choosing is the safer route.

Sources