Aleph Alpha 78B MoE Kolibri-1 Community GGUF Released: Reaching 13-15 TPS on CPU

Community GGUF quantizations for Aleph Alpha's 78B MoE model Kolibri-1 are now available, achieving 13-15 TPS on a Ryzen 7 7800X3D CPU using 128GB RAM and patch

tau · October 5, 2026

#Kolibri #AlephAlpha #GGUF #llamacpp #LocalLLM #MoE #OpenSource

Aleph Alpha 78B MoE Kolibri-1 Community GGUF Released: Reaching 13-15 TPS on CPU

Community GGUF quantizations for Kolibri-1, a 78-billion-parameter Mixture-of-Experts (MoE) reasoning model released under the Apache 2.0 license by German AI company Aleph Alpha, have officially arrived on Hugging Face. Created by open-source developer Hob-forge following the official weight drop, this community release demonstrates that high-parameter reasoning models can achieve usable inference throughput on standard consumer CPUs without requiring enterprise-grade accelerator hardware.

Terminal benchmark output and hardware specifications for Aleph Alpha Kolibri-1 78B MoE GGUF running on an AMD Ryzen 7 7800X3D CPU

Image source: @TeksEdge / Aleph Alpha

While Kolibri-1 features 78 billion total parameters (78,103,074,560), its sparse MoE architecture dynamically activates only about 3.46 billion parameters (3,457,573,120) per token. By routing tokens through 6 selected experts out of 384 along with 1 shared expert, the model drastically reduces per-token arithmetic demands, making real-time text generation feasible on consumer CPUs as long as host system memory can hold the model weights.

Kolibri-1 Architecture and Community GGUF Quantization

Aleph Alpha officially launched Kolibri-1 on October 3, 2026, targeting European enterprise and public administration use cases with a bilingual focus on German and English. The model incorporates a long-context window of up to 1,048,576 tokens (with recommended serving at or below 262,144 tokens for optimal efficiency), explicit reasoning modes, arithmetic capabilities, and structured tool calling.

The vendor's official release checkpoint shipped in an FP8 format (float8_e4m3fn with 128×128 block scales), weighing roughly 78GB—a footprint that cannot be loaded onto any single consumer GPU. In response, Hob-forge converted the FP8 checkpoint by first dequantizing the weights into BF16 format and then directly quantizing them into GGUF containers. Importantly, this process is not a lossy requantization from an existing low-bit release, preserving weight fidelity.

  • Q4_K_M GGUF (Kolibri-1-Q4_K_M.gguf): Sized at approximately 47.5GB as a single file, serving as the primary recommended quantization for systems with sufficient host RAM.
  • Q8_0 GGUF (Kolibri-1-Q8_0): Sized at roughly 83.1GB total, distributed in two split parts to accommodate file system limits. Pointing llama.cpp at the first part automatically loads the second.

CPU-Only 13-15 TPS Benchmark and 128GB RAM Requirements

According to benchmark results shared by tester David Hendrickson (@TeksEdge), Hob-forge evaluated the Q4_K_M conversion purely on an AMD Ryzen 7 7800X3D desktop CPU paired with 128GB of system RAM, operating without any discrete GPU acceleration. The setup achieved decoding speeds of approximately 13 to 15 tokens per second (TPS).

The evaluation verified reliable performance across German and English generation tasks, arithmetic word problems (with reasoning enabled), and tool calling operations. Because only 3.46B active parameters are evaluated per token, memory bandwidth pressure during autoregressive decoding is significantly lower than that of conventional dense networks in the same parameter tier.

However, system memory headroom remains a strict constraint. With the Q4_K_M binary alone occupying 47.5GB, running on a 64GB machine leaves very little buffer for OS overhead and KV caches, creating a substantial risk of memory paging or out-of-memory (OOM) crashes. As a result, 128GB of system RAM is practically required for uninterrupted local CPU operation.

Patched llama.cpp Requirement and Hardware Acceleration Caveats

Developers planning to test Kolibri-1 locally should note several immediate operational caveats:

  • Patched build requirement: Because Kolibri-1 uses a distinct MoE routing topology, running the model currently requires a custom patched build of llama.cpp. Support for the kolibri1 architecture has not yet been merged into upstream llama.cpp.
  • Hardware acceleration untested: The initial release and published throughput figures focus strictly on CPU inference. The author explicitly notes that GPU acceleration pathways, including CUDA, Vulkan, and Metal backends, have not yet been fully validated.
  • Packaged runtime lag: Turnkey local runners like Ollama or LM Studio will only support the model once the architectural patch is merged into upstream llama.cpp and propagated into their respective releases.

By bringing a sovereign 78B European MoE model into practical desktop CPU reach within days of its open-weight release, this community quantization marks a notable milestone for local inference ergonomics and developer experimentation.

Sources