Nari Labs Launches 40ms Ultra-Low-Latency Streaming STT Endpoint Powered by Qwen3-ASR 1.7B

Nari Labs unveiled an ultra-low-latency streaming STT endpoint powered by Qwen3-ASR 1.7B and a custom inference engine, delivering 40ms TTFS at $0.06/hr, now in

tau · September 12, 2026

#NariLabs #STT #Qwen3ASR #SpeechRecognition #LowLatency #AIEndpoint #OpenSourceAI

Nari Labs Launches 40ms Ultra-Low-Latency Streaming STT Endpoint Powered by Qwen3-ASR 1.7B

On September 12, 2026 (KST), Y Combinator-backed voice AI startup Nari Labs officially launched an ultra-low-latency streaming Speech-to-Text (STT) endpoint powered by the open-source Qwen3-ASR 1.7B model paired with a proprietary inference engine. The service is currently accessible to developers and teams in an open public beta directly through its official website at narilabs.com.

Nari Labs streaming speech-to-text live transcription demo and 40ms latency benchmark metrics powered by Qwen3-ASR 1.7B

Image source: Toby Kim (@doyeob) via X

As demand surges for conversational voice agents, live transcription, and real-time audio workflows, response latency and infrastructure costs have remained persistent roadblocks. Nari Labs aims to address both hurdles simultaneously by coupling a lean open-source foundation model with a custom-engineered streaming inference pipeline.

40ms Ultra-Low Latency via Qwen3-ASR 1.7B and Custom Inference

The technological centerpiece of Nari Labs' new endpoint is the integration of the open-source Qwen3-ASR 1.7B speech recognition model with an in-house inference engine optimized specifically for chunk-by-chunk audio streaming.

According to performance measurements cited by Nari Labs from the independent voice AI benchmarking platform Coval (recorded on September 10, 2026), the endpoint established standout rankings across two primary evaluation criteria:

  • Industry-leading latency: The system registered a median (p50) latency of 43ms and a Time to First Segment (TTFS) of 40ms, ranking first overall in speed across all benchmarked speech-to-text models.
  • Competitive transcription accuracy: It achieved a Word Error Rate (WER) of 3.3%, securing third place among all publicly available speech recognition APIs evaluated on the platform.

Unlike traditional large-scale voice architectures that incur hundreds of milliseconds of buffering delay to process full utterance contexts, Nari Labs utilizes a compact 1.7-billion-parameter model footprint coupled with an aggressive streaming pipeline designed to return transcribed segments almost synchronously with speech input.

$0.06/Hour Pricing and Competitive Efficiency Benchmarks

Alongside rapid response times, Nari Labs has placed aggressive pricing at the forefront of its market strategy.

The company set its standard streaming STT service rate at $0.06 per audio hour ($0.06/hr). When evaluated against existing commercial speech offerings in production, Nari Labs reported significant advantages in both throughput and operating expenditures:

  • Throughput versus ElevenLabs: The endpoint delivers approximately three times the processing speed of ElevenLabs Scribe V2 in real-time conversational streaming tests.
  • Cost reduction versus Google Gemini: It offers an operating cost structure approximately nine times cheaper than Google's Gemini Transcribe 3.5 API, substantially reducing the financial burden for high-volume audio processing pipelines.

While currently available at no charge during its public beta period, Nari Labs confirmed that the $0.06/hr rate will serve as the permanent baseline fee once the service transitions into commercial availability.

Public Beta Deployment and Community Observations

Nari Labs' ultra-low-latency endpoint has drawn immediate attention across latency-critical domains, including interactive AI voice agents, live captioning systems, customer service call centers, and in-game communication channels.

Nevertheless, beyond promising synthetic benchmark results, initial hands-on testing by developers has surfaced several key operational observations and areas for continued refinement:

  • Multilingual performance variation: While English transcription demonstrates high fidelity, early users observed notable quality degradation and occasional dropped phrases when processing non-English speech, such as Japanese audio streams.
  • Web playground traffic bottlenecks: Surging traffic following the initial launch triggered intermittent buffering delays and temporary throughput throttling on the browser-based testing playground.
  • Real-world acoustic validation: Because the highlighted 40ms TTFS and 3.3% WER metrics originate from Coval's controlled testing environment, ongoing developer validation will be necessary across diverse background noise profiles, accents, and varying network conditions.

Nari Labs stated that it is actively incorporating developer feedback during the public beta to stabilize multilingual transcription and expand its dedicated serving infrastructure.

Sources