Mitsuba & HiMitsuba 27B GGUF: A 1.58-Bit Vision-Language Model for ComfyUI on 16GB GPUs

Mitsuba and its uncensored HiMitsuba LoRA compress Qwen3.8-27B into a 7.3GB 1.58-bit VLM, enabling fast local image analysis and prompt generation for ComfyUI o

tau · October 6, 2026

#ComfyUI #GGUF #VLM #Qwen3.8 #llamacpp #LocalAI

Mitsuba & HiMitsuba 27B GGUF: A 1.58-Bit Vision-Language Model for ComfyUI on 16GB GPUs

Open-source AI developer isichan (isichan-ai) has released the 'Mitsuba & HiMitsuba 27B GGUF' model suite, compressing the 27-billion-parameter Qwen3.8-27B vision-language model into a 1.58-bit ternary quantization that runs fully inside the VRAM budget of a single 16GB GPU. Graphic representation of Mitsuba and HiMitsuba 27B GGUF vision-language model architecture and ComfyUI prompt generation benchmarks on a 16GB GPU

Image source: @HuggingModels

Engineered specifically for local creative pipelines such as ComfyUI and Krea 2, this vision-language model (VLM) analyzes input reference images and synthesizes highly detailed prompts adhering to strict downstream generation constraints. On October 4, 2026, the repository was officially renamed from Mitsuba-ComfyUI-27B-GGUF to include the newly added uncensored HiMitsuba LoRA weights, significantly mitigating safety refusals during creative image and video generation.

1.58-Bit Ternary Quantization and Single 16GB GPU Local Inference

The primary architectural breakthrough of Mitsuba 27B is packing a capable 27B parameter multimodal model into consumer-tier GPU hardware through ternary quantization.

  • PQ2_0 Format (7.32 GB): With an on-disk footprint of 7.32 GB, this build fits comfortably within 16GB VRAM alongside active inference buffers, eliminating memory pressure.
  • PTQ1_0 Format (6.00 GB): An ultra-compact 5.998 GB quantization variant tailored for severely memory-constrained local host environments.
  • Multimodal Projector Requirement: Processing vision inputs requires loading a dedicated 0.63 GB multimodal vision projector file (mmproj-Q8_0.gguf) alongside the base weights.
  • Decoding Performance: Measured on an RTX 5090 GPU, standalone PQ2_0 achieves a decoding throughput of approximately 119 tokens per second.

Despite aggressive ternary compression, vision benchmark performance drops only slightly from 89.8 to 87.8 compared to unquantized baselines, while prompt instruction-following adherence noticeably improves over standard quantized alternatives.

ComfyUI-Specialized Vision Prompting and Strict Non-Coding Boundary

Unlike general-purpose conversational models, Mitsuba is intentionally tuned exclusively for image analysis, captioning, and prompt synthesis in creative workflows.

  • Image Analysis and Reverse Prompting: Given a visual input, the model breaks down composition, subject details, lighting, perspective, and styling to generate structured prompts aligned with specified system directives.
  • Video Motion Descriptions: Extends beyond static single-image captions to produce consistent motion descriptions required for generative video pipelines.
  • Strict Non-Coding Boundary: Model weights have been specialized solely for visual description and prompt syntax. Consequently, general coding capability collapses to 4/100, making Mitsuba entirely unsuitable for software development or programming assistance.

Uncensored HiMitsuba LoRA Integration and Refusal Reduction

The accompanying HiMitsuba-Uncensored-LoRA.gguf adapter (0.07 GB) addresses conservative safety barriers that frequently cause vision models to decline descriptive art requests.

  • Dramatic Refusal Drop: On a benchmark of 106 test prompts, outright refusals decreased from 51 to just 7.
  • Evasion Mitigation: Evasive responses—where the model replies while stripping critical details—fell from 73% to 32% on visual description tasks, reducing combined failures from 59 to 13 out of 106.
  • Uncensored Metric Surge: The overall uncensored benchmark score climbed from 53.2 to 89.6.
  • Prompt Adherence Gains: Strict condition-following prompt generation improved from 6/10 to 8/10 with the LoRA applied.
  • PQ2_0 Exclusivity: The HiMitsuba LoRA is calibrated exclusively for the 7.32 GB PQ2_0 build and is not supported on the PTQ1_0 variant.
  • Performance Tradeoffs: Applying the LoRA at runtime slightly reduces RTX 5090 throughput from ~119 t/s to ~103.1 t/s and lowers tool-calling persistence (self-restraint drops from 62.7 to 48.2, aborting after a single tool error). VRAM consumption remains unaffected.

PrismML llama.cpp Fork Runtime and Mandatory Disabled Reasoning

Deploying Mitsuba 27B requires specific runtime considerations due to its unique ternary quantization structure.

  • PrismML llama.cpp Fork: Because upstream llama.cpp does not yet support ternary PQ2_0 and PTQ1_0 quantization types, users must run the prism branch from the PrismML fork.
  • Mandatory Disabled Reasoning: Mitsuba was trained strictly with reasoning turned off. Enabling thinking mode causes the model to loop identical sentences inside its internal chain-of-thought buffer without emitting a final response.
    • Server CLI flag: --reasoning off
    • Request payload: "chat_template_kwargs": {"enable_thinking": false}

The verified deployment command tested for evaluation is:

llama-server -m Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf \
  --mmproj mmproj-Q8_0.gguf \
  --no-mmproj-offload \
  --reasoning off \
  --jinja \
  --temperature 0.6 \
  --top-k 20 \
  --top-p 0.95 \
  --ctx-size 131072 \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --n-gpu-layers 99

For digital artists and ComfyUI developers seeking local multimodal capabilities on consumer hardware, Mitsuba & HiMitsuba 27B offers an efficient and practical solution.

Sources