Qwen3.5-0.8B Japanese Ultralight Model v2 Released: Smoother Japanese, Weights Available Now

Developer atemoya has released v2 of a Qwen3.5-0.8B-based Japanese SFT model on Hugging Face. The update focuses on more natural Japanese output rather than kno

tau · October 3, 2026

#Qwen3.5 #OnDeviceAI #HuggingFace #SFT #SmallLLM

Qwen3.5-0.8B Japanese Ultralight Model v2 Released: Smoother Japanese, Weights Available Now

Developer atemoya (Hugging Face account Takenoko12345678) released Qwen3.5-0.8B-Japanese-SFT-v2 on Hugging Face on October 3, 2026. It is the second generation of a Japanese-specialized SFT model built on Qwen3.5-0.8B. A formal announcement will follow later, but according to the author's own post, the weights are already downloadable.

Concept visual for Qwen3.5-0.8B ultralight Japanese SFT model v2 release on Hugging Face

Image source: https://huggingface.co/Takenoko12345678/Qwen3.5-0.8B-Japanese-SFT-v2

This v2 keeps the 0.8B-parameter ultralight architecture and focuses on improving the naturalness and fluency of Japanese output rather than expanding the sheer amount of knowledge, according to the author. A browser-based web inference environment is still in preparation, so for now users need to download the weights from Hugging Face and run the model locally.

What changed in v2 and how to use it now

The author's direct description of v2 comes down to two points.

  • More natural Japanese: Knowledge has "not grown all that dramatically," but Japanese naturalness should be considerably better, in the author's words.
  • Access path: Weight downloads from the Hugging Face repository are available immediately, while web inference is still in preparation.

As an ultralight model, local on-device execution is the premise, continuing the direction set out at the v1 release for PCs without a discrete GPU.

Confirmed v1 baseline (do not confuse with v2 figures)

The following figures come from the author's October 2 post about v1, released a day earlier. They are not independent benchmark figures for v2.

  • Base model: A Japanese model starting from Qwen3.5-0.8B.
  • Author-stated v1 baseline: Japanese commonsense question accuracy rose from 41.7% to 69.4%, with responses returning in about one second per turn on a PC without a GPU.
  • Training data: The author stated the training data was fully published at the v1 release.
  • v1 browser inference: A separate browser inference link was provided for v1, with a note that it does not work in private browsing mode.

In short, v2 is a refinement release that keeps v1's lightweight, local-execution approach while polishing expression quality. Numerical comparisons will have to wait until v2-specific measurements appear.

Sources