TAU-HOME.COM
LOADING

Upscaling to 8K With LLM Vision Instead of a Dedicated Upscaler

A patch-upscale-then-stitch LLM upscaling workflow that uses a frontier model's vision and scene understanding to reach 8K, plus its time cost and limits.

tau · October 8, 2026

#AIimages #upscale #ChatGPT

Upscaling to 8K With LLM Vision Instead of a Dedicated Upscaler

This is a write-up of the LLM upscaling workflow that original author @flowersslop (Flowers) shared on X on October 7, 2026. After viewers assumed a dedicated upscaler had been used, the author shared the actual approach as a reusable skill.

Close-up detail comparison showing AI-upscaled image texture reconstructed with scene understanding

Image source: @flowersslop / X

The core idea is simple. Instead of enlarging pixels, throw frontier-model vision and world knowledge at the upscale so the scene gets reconstructed. According to the author, dedicated upscalers are decent at textures and local detail but start to fall apart once the image needs actual scene understanding, semantics, or visual reasoning.

Workflow: upscale every patch, then stitch into one 8K image

The workflow the author described:

  • Model used: ChatGPT 6.1 sol.
  • Method: it does not upscale individual patches in isolation; it upscales all patches and stitches them back together into one big 8K image end to end.
  • Possible alternative: Nano Banana via API may work even better. The author added a personal preference for Nano Banana over ChatGPT images.

The full skill text is not reproduced here. The evidenced source text is truncated mid-sentence ("paste this into codex o…"), so no commands or prompts beyond the evidence are quoted or reconstructed.

Time cost and how to read the comparison

  • Speed: around 10 minutes for two 8K passes. Slowness is the stated catch.
  • Comparison images: the X comparison undersells the result, in the author's words, because X/Twitter compression destroys fine detail.

Practically, this is an approach worth trying when reconstructed texture matters more than pixel-exact fidelity. Budget for the 10-minute time cost, and note the limit that the complete skill/prompt body is not fully present in the public evidence.

Original source