Ollama Cloud Rolls Out DeepSeek-V4.1-Flash with US-EU Hosting and Zero Data Retention
Ollama Cloud has deployed DeepSeek-V4.1-Flash with US-EU hosting, Zero Data Retention, official API price matching including 50% off-peak discounts, and pay-as-
On September 11, 2026, Ollama announced via its official X account (@ollama) that DeepSeek's latest multimodal flagship model, DeepSeek-V4.1-Flash, has been fully rolled out across the Ollama Cloud infrastructure. This release allows software developers and enterprise engineering teams to run large-scale Mixture-of-Experts (MoE) inference directly in the cloud without being limited by local GPU hardware constraints.

Image source: @ollama (X) / Ollama Library
Rather than a simple catalog addition, the rollout introduces an enterprise-grade Zero Data Retention policy, multi-region distributed hosting across the United States and Europe, exact token-level price matching with DeepSeek's official API, and linked 50% off-peak pricing discounts.
Distributed US and European Hosting with Zero Data Retention
When adopting cloud-based AI inference in production environments, development teams frequently face challenges around data privacy, regulatory compliance, and network latency. Ollama Cloud addresses these governance and operational requirements by establishing a hardened hosting setup.
- Distributed US and European Regions: By deploying inference endpoints across geographically separated data centers in North America and Europe, the platform delivers reduced latency and resilient request routing for international workloads.
- Strict Zero Data Retention Policy: Neither incoming user prompts nor generated model completions are logged or retained on Ollama Cloud servers. Furthermore, user inputs and outputs are never stored or used to train, fine-tune, or evaluate future AI models.
This privacy-first posture removes a major obstacle for enterprise organizations and privacy-sensitive engineering teams that previously avoided external hosted LLMs due to data exposure concerns.
552B MoE Multimodal Architecture and Cloud Inference Requirements
The DeepSeek-V4.1-Flash release deployed on Ollama Cloud features an architecture designed to combine high-level reasoning with computational efficiency. Operating on a 552B total parameter Mixture of Experts (MoE) design, the model achieves rapid response times while preserving deep domain competence.
- Up to 1 Million (1M) Token Context Window: The model supports long-context analysis capable of digesting complete codebases, technical reference manuals, and extensive multi-document sets within a single inference session.
- Native Visual Understanding: Moving beyond text processing, the model handles multimodal inputs including architectural diagrams, charts, UI screenshots, and complex images directly within the conversation thread.
- Speed and Efficiency Gains: According to official performance metrics, DeepSeek-V4.1-Flash delivers inference throughput and operational cost-efficiency that surpass the earlier DeepSeek-V4-Pro.
Crucially, unlike standard Ollama models that run locally on on-premise hardware, this model operates via Ollama Cloud under the deepseek-v4.1-flash:cloud identifier. Running inference requires an active internet connection and authentication with an Ollama account.
Official API Price Parity, 50% Off-Peak Discounts, and Flexible Billing
Ollama Cloud adopts a direct pass-through pricing model without the markups or platform margins typically seen in managed API hosting providers.
- Official API Price Matching: Token consumption is priced at the exact per-token rates set by DeepSeek's official API, letting developers build on Ollama's unified interface without incurring a pricing penalty.
- 50% Off-Peak Discount Integration: Ollama Cloud directly mirrors DeepSeek's off-peak scheduling, applying a 50% discount to token rates during designated low-traffic hours.
- Pay-as-You-Go Access for Free Accounts: Users can start making API calls immediately on a pay-as-you-go basis without monthly lock-ins, and standard access is also included for subscribers of Ollama Pro, Max, and Team plans.
Because token rates differ substantially between standard and off-peak windows, engineering teams executing high-volume batch processing or automated pipelines can schedule workloads during off-peak hours to cut operational expenses in half.
Sources
- Ollama Library: Ollama Library - DeepSeek-V4.1-Flash
- Ollama Official X (@ollama): DeepSeek-V4.1-Flash Full Rollout Announcement