P-Video-2

P-Video-2 is Pruna’s premium video generation model, the quality-focused successor to P-Video. Generate from text, an image, or audio-conditioned input in one endpoint. Compared with P-Video, the quality jump is native-speech lip-sync (pause / line / rest), sharper close-ups, and identity lock.

Designed for production results, it offers text-to-video, image-to-video, and audio-conditioned generation at up to 1080p and 48 fps, with optional duration up to 20 seconds (or leave duration empty and the model chooses length from the prompt). Use draft mode for faster, lower-cost iteration, then switch to full quality for finals.

Note

When using P-Video-2, make sure to respect the copyright of the images and audio you use as input and the content you generate.

Pricing: billed per second of the video returned (measured from the finished clip in whole seconds).

Resolution

Draft mode is OFF

Draft mode is ON

720p

$0.025 per second of output video

$0.015 per second of output video

1080p

$0.05 per second of output video

$0.03 per second of output video

Speed: model execution time per second of output video.

Resolution

Draft mode is OFF

Draft mode is ON

720p

~0.91 s per second of output video

~0.41 s per second of output video

Tip

Test it now in the P-Video-2 Playground.

Prompt formula

Fast pass

One prompt: subject, action, scene. Enough for first looks. Iterate in draft.

[prompt]
subject, action, scene

Locked-in

Add camera, lighting, style, and audio. Optional still or imported track for repeatable runs.

[prompt] [image] [audio]
subject, action, scene + camera, lighting, style, dialogue, optional still for I2V, native speech or imported track
Fast pass: text-to-video; subject, action, and scene only.
Locked-in: text-to-video; camera, lighting, style, and audio spelled out.
  1. prompt: required. Name the subject, action, and scene. Add camera, lighting, style, and audio when you need a repeatable final.

    • Fast pass: “Knitted purple prune character, reading a book, cozy room.”

    • Locked-in: “Knitted purple prune character, reading a book, eyes look at book then at camera, cozy room, camera slowly zooming in, warm natural light, documentary style. Character reading aloud: ‘P-Video-2 generates premium AI video with sharp lip-sync and native audio.’”

  2. image: optional still for image-to-video (jpg, jpeg, png, webp). Keep motion subtle and aligned with the reference frame. aspect_ratio is ignored when an image is set. Optional last_frame_image for end-frame control.

    • Fast pass: omit for text-to-video.

    • Locked-in: a high-quality, well-lit still; generate it with P-Image-Ideogram when you need a locked composition.

  3. audio: optional imported track (flac, mp3, wav). Duration follows the audio and duration is ignored. Lead lip-sync demos with native speech (save_audio: true, no imported audio).

    • Fast pass: save_audio: true; name dialogue or music in the prompt.

    • Locked-in: paste full lyrics when singing; import audio only when you need a specific track.

Slot

Fast pass (enough to run)

Locked-in (stronger control)

prompt

Subject, action, scene.

Subject + action + scene + camera, lighting, style, audio; same wording every rerun.

image

Empty (text-to-video).

Well-lit still; optional last_frame_image.

audio

Native speech in the prompt; save_audio: true.

Imported track when timing must match; otherwise keep native speech for lip-sync.

Tip

For comprehensive video prompting (motion, framing, atmosphere), see the Video Generation guide.

Choosing the right video model

Pruna ships performance video models that share the same prediction API, but each solves a different production problem. P-Video-2 generates new footage (Model: p-video-2).

P-Video-2-Pro

P-Video-2

P-Video

P-Video-Avatar

P-Video-Animate

P-Video-Replace

P-Video-Edit

One-line job

Generate cinematic footage from prompts

Generate premium footage from prompts

Generate fast / affordable footage

Speak from one still (script or audio)

Retarget one still with clip motion

Swap characters in existing footage

Rewrite content in existing footage

You start with

Text prompt (+ optional first / last frame)

Text prompt (+ optional image / audio)

Text prompt (+ optional image / audio)

Portrait still + voice_script or audio

Source video + one still

Source video + identity stills

Source video + text prompt

You keep from the source

N/A (new scene); canvas follows reference images when set

N/A (new scene); strong I2V consistency when an image is set

N/A (new scene)

Aspect ratio of the still

Motion, timing, camera from the driver

Camera, timing, blocking, background

Camera, timing, subject performance

Typical ask

“Make an 8 s cinematic product or documentary clip with generated audio.”

“Make a 10 s product ad with music and lip-sync.”

“Draft a 5 s social clip cheaply and fast.”

“This spokesperson says this line in French.”

“Animate this catalog still using our winning ad take.”

“Put our creator in this UGC b-roll.”

“Turn this silver SUV red and keep the camera move.”

Quick decision guide

  • Need the highest generation quality, lip-sync, or audio-conditioned output → P-Video-2.

  • Need the fastest / cheapest generation iteration → P-Video.

  • Single-speaker talking head, voice only, no music bed → P-Video-Avatar.

  • Footage exists and the hero still should move like the driver → P-Video-Animate.

  • Footage exists and you need different people in the same shot → P-Video-Replace.

  • Footage exists and you need a color, product, environment, object, or text edit → P-Video-Edit.

Key features

Marketing and ads

Text-to-video social ads, image-to-video from brand stills, and clips with music or SFX. Prefer P-Video-Avatar only for voice-only talking heads.

Media and entertainment

Music visuals, short-form entertainment, and audio-driven performance clips. Paste full lyrics when singing; duration follows imported audio.

Retail and e-commerce

Product loops and shoppable lifestyle clips. Keep motion subtle and start from a clean product or on-model still.

Corporate and education

Training, explainers, and comms that need a music bed, ambience, or two speakers. Voice-only single-presenter scripts → P-Video-Avatar.

Gaming

Trailers, world reveals, and character moments. One hero action per clip; prefer 24 fps at 1080p.

Native-speech lip-sync

When the model generates the voice (save_audio: true, no imported audio), it pauses, delivers the line, then rests, and the mouth follows.

Image-to-video + last frame

Animate a reference still; optional last_frame_image for end-frame control.

Draft mode

Faster, lower-quality preview at half the per-second rate. Iterate with draft: true, then finalize with draft: false.

Practical constraints

  • Maximum duration is 20 seconds.

  • When audio is provided, output length and billing follow the audio; duration is ignored.

  • aspect_ratio is ignored when an input image is provided.

  • Input images: jpg, jpeg, png, webp. Input audio: flac, mp3, wav.

  • Draft mode is a preview: faster and cheaper, lower quality.

  • Above two speakers, speaker separation can degrade.

Examples

Integration

P-Video-2 uses the same Pruna predictions API as other performance models. Text-to-video needs only a prompt; image-to-video and audio-conditioned runs upload files first.

Tip

For more information on how to use the API, see the API Reference.

API endpoint

Base URL: https://api.pruna.ai/v1/predictions

Authentication

-H 'apikey: YOUR_API_KEY'
-H 'Model: p-video-2'

Text-to-video (asynchronous)

curl -X POST 'https://api.pruna.ai/v1/predictions' \
  -H 'Content-Type: application/json' \
  -H 'apikey: YOUR_API_KEY' \
  -H 'Model: p-video-2' \
  -d '{
    "input": {
      "prompt": "A sports car drifting through a neon-lit city at night, cinematic aerial shot",
      "duration": 5,
      "resolution": "720p",
      "aspect_ratio": "16:9"
    }
  }'

Text-to-video (synchronous)

curl -X POST 'https://api.pruna.ai/v1/predictions' \
  -H 'Content-Type: application/json' \
  -H 'apikey: YOUR_API_KEY' \
  -H 'Model: p-video-2' \
  -H 'Try-Sync: true' \
  -d '{
    "input": {
      "prompt": "A sports car drifting through a neon-lit city at night, cinematic aerial shot",
      "duration": 5,
      "resolution": "720p",
      "draft": true
    }
  }'

Image-to-video

Upload a reference image, then pass its file URL:

curl -X POST "https://api.pruna.ai/v1/files" \
  -H "apikey: YOUR_API_KEY" \
  -F "content=@/path/to/your/file.jpg"
curl -X POST 'https://api.pruna.ai/v1/predictions' \
  -H 'Content-Type: application/json' \
  -H 'apikey: YOUR_API_KEY' \
  -H 'Model: p-video-2' \
  -d '{
    "input": {
      "prompt": "The camera slowly pushes in, the person turns their head and smiles",
      "image": "https://api.pruna.ai/v1/files/fqadqq42xq",
      "duration": 5,
      "resolution": "720p"
    }
  }'

Configuration

Required parameters

Parameter

Type

Description

prompt

string

Text prompt for video generation

Optional parameters

Parameter

Type

Default

Description

duration

integer

—

Duration in seconds (1–20). Leave empty to let the model choose from the prompt. Ignored when audio is provided

image

string (URI)

—

Input image for image-to-video (jpg, jpeg, png, webp)

audio

string (URI)

—

Input audio to condition generation (flac, mp3, wav)

resolution

string

"720p"

"720p" or "1080p"

fps

integer

24

24 or 48

aspect_ratio

string

"16:9"

"16:9", "9:16", "4:3", "3:4", "3:2", "2:3", "1:1". Ignored when an input image is provided

seed

integer

random

Random seed for reproducible generation

draft

boolean

false

Faster, lower-quality preview. Billed at half the standard per-second rate

save_audio

boolean

true

Save the video with audio

last_frame_image

string (URI)

—

Reference image for the last frame

prompt_upsampling

boolean

true

Enhance the prompt

disable_safety_filter

boolean

true

Disable safety filter for prompts and input images

Argument recommendations

Use these patterns for consistent quality:

  • prompt: follow subject / action / scene, then optional camera, lighting, style, and audio. For image-to-video, keep motion subtle and aligned with the reference frame.

  • image: high-quality, well-lit still; aspect_ratio is ignored when an image is set. Optional last_frame_image for end-frame control.

  • audio: use imported audio when you need a specific track; duration follows the audio and duration is ignored. Lead lip-sync demos with native speech (save_audio: true, no imported audio).

  • duration: 1–20 seconds, or leave empty so the model chooses from the prompt.

  • resolution / fps: iterate in 720p with draft: true, then rerun finals with draft: false. Prefer 24 fps at 1080p for stability; use 48 fps at 720p for smoother action.

  • prompt_upsampling: leave true for production; compare true vs false on the same seed before scaling.