Strata v0.1.34 Released: AMD GPU Acceleration on Windows and Built-in MCP Server Support

Open-source local LLM engine Strata v0.1.34 adds official AMD GPU acceleration on Windows and a built-in MCP server for direct AI coding agent integration.

tau · October 3, 2026

#Strata #AMD #Windows #MCP #LocalLLM #DevTools

Strata v0.1.34 Released: AMD GPU Acceleration on Windows and Built-in MCP Server Support

Developer Niko Veit (@coldniko) has released v0.1.34 of Strata, an open-source local Large Language Model (LLM) inference engine. This update introduces official AMD GPU hardware acceleration on Windows alongside a built-in Model Context Protocol (MCP) server designed to connect directly with external AI agent ecosystems.

Strata local LLM inference engine supporting AMD GPU acceleration on Windows and built-in MCP server

Image source: Niko Veit (@coldniko) / Niko1221 Strata

Official AMD GPU Acceleration on Windows and Hardware Reach

Expanding beyond NVIDIA CUDA-dominated local inference setups, Windows users with AMD hardware can now utilize Strata with direct GPU acceleration.

  • Reported Radeon Pro Runs: One community user reported running an IQ2 quantized model with a 256K context on an AMD Radeon Pro W7900 system. The developer replied that another user is running a W7800 with 48GB of RAM. These are individual user reports, not representative performance measurements.
  • Initial Windows AMD Phase: Windows AMD support is in its initial release phase; the developer's release note says issue fixes are ongoing alongside the release.

Built-in MCP Server for Native AI Agent Integration

Strata v0.1.34 includes a Model Context Protocol (MCP) server feature.

  • Direct Agent Interoperability: External AI agents can discover and communicate with Strata directly, without intermediate bridges or proxy layers.
  • Local Infrastructure for Agent Workflows: The feature targets wiring local compute resources directly into agent workflows.

Ongoing Roadmap and Architectural Considerations

The Strata project continues to expand performance optimizations and hardware compatibility.

  • KV Cache Optimizations in Development: After a user asked for the KV cache optimization techniques implemented in FreeTokens, the developer replied that the work is in progress.
  • Mixed GPU Configurations and Legacy Architectures: Split execution across mixed CUDA and HIP devices and support for older GPU architectures (Pascal, gfx1031, and similar) were raised as community questions; the acquired evidence contains no confirmed developer answer.

Sources