Pruna AI Launches P-Video-2: Multimodal Video Model with Native Audio and 48fps Output

Pruna AI has released P-Video-2, supporting up to 1080p at 48fps, 20-second clips, native 48kHz audio, and inference from $0.015/sec across multiple cloud provi

tau · September 12, 2026

#AI #VideoGeneration #P-Video-2 #PrunaAI #Multimodal #Replicate #Segmind

Pruna AI Launches P-Video-2: Multimodal Video Model with Native Audio and 48fps Output

On September 10, 2026, artificial intelligence optimization and efficiency provider Pruna AI officially announced P-Video-2, a next-generation multimodal video generation foundation model capable of processing text, image, and audio inputs through a single unified endpoint. Moving past legacy workflows that generate visuals and audio in disconnected stages, the new model synthesizes high-fidelity video alongside synchronized 48kHz stereo sound within a single integrated inference pass.

Official brand visual artwork for Pruna AI's P-Video-2 multimodal video generation model

Image source: Pruna AI (playground.pruna.ai)

The launch is aimed squarely at the demands of modern production environments, addressing the need for higher resolutions, flexible framerate control, near-real-time generation speeds, and predictable compute costs. By providing granular rendering modes, P-Video-2 enables creators and software developers to handle rapid storyboarding, visual exploration, and production mastering within an unbroken pipeline.

Unified Multimodal Inputs and Synchronized 48kHz Stereo Audio Generation

The architectural centerpiece of P-Video-2 is its native integration across input modalities alongside direct simultaneous audio synthesis.

Traditional generative video pipelines have largely separated text-to-video (T2V) and image-to-video (I2V) workflows, frequently relying on secondary audio models to stitch background soundscapes or Foley effects in post-production. Pruna AI consolidates these disparate steps into a singular API structure.

  • Unified Conditioning Inputs: Creators can submit standard text prompts, reference still images, or audio tracks as conditioning inputs through the unified endpoint.
  • Synchronized 48kHz Stereo Audio: High-fidelity stereo audio is synthesized in lockstep within the visual generation pass, producing an integrated, playback-ready video asset without requiring downstream sound design passes.
  • Audio-Driven Timeline Alignment: When audio conditioning is provided, the total duration of the generated video aligns automatically to the uploaded audio track, bypassing generic API duration parameters.

This unified interface significantly reduces iteration friction for short narrative vignettes, music-synchronized videos, motion graphics, and high-tempo social clips.

Specifications up to 1080p 48fps with Draft and Standard Pricing Tiers

In terms of visual output quality and operational flexibility, P-Video-2 introduces a balanced performance profile engineered for production deployment.

The model supports resolutions up to 1080p and extends beyond standard 24fps cinema playback to offer smooth 48fps high-frame-rate rendering. Single generation runs can yield coherent clips spanning up to 20 seconds in length.

To accommodate different stages of the creative process, Pruna AI offers two distinct rendering tiers:

  • Draft Mode: Engineered for swift ideation and rapid storyboard iteration. It renders at approximately 0.41 seconds of compute per output second, with pricing starting at $0.015 per second for 720p output.
  • Standard Mode: Tailored for final delivery and production-grade rendering fidelity. It processes at roughly 0.91 seconds per output second, priced from $0.025 per second at 720p.

This granular per-second pricing structure allows engineering teams and studio pipelines to forecast operational costs accurately across large batch operations and interactive services.

Cloud Deployment Ecosystem and Upcoming Partner Benchmarks

Rather than restricting access to a proprietary walled garden, Pruna AI made P-Video-2 available on day one across an expansive network of third-party cloud inference providers.

Developers and enterprise teams can immediately deploy and invoke the model within their existing infrastructure environments. Live partner availability includes:

  • Supported Cloud Platforms: Replicate, Segmind, Cloudflare, Runware, RunPod, Eachlabs, Wavespeed AI, and Eden AI.

This multi-provider distribution ensures broad availability, low latency across regions, and zero vendor lock-in for organizations integrating video generation into their commercial applications.

Regarding objective performance metrics, Pruna AI noted that detailed benchmark results conducted through official evaluation partners—including Datapoint AI, Rapidata, and DesignArena—will be published in the period following the launch. This third-party validation approach aims to establish transparent, reproducible benchmarks for image coherence, motion stability, and multimodal alignment.

Sources