# Kling AI Avatar v2 Pro

> Talking or singing avatar video from one portrait and one audio track. The face lip-syncs to your speech or song, your audio plays through unchanged, and the clip runs as long as the audio (5 to 60 s). Send image_url, audio_url (PCM WAV only for now) and duration_s (your audio's length); billed per second of the finished clip.

Human version: https://kymaapi.com/models/kling-avatar-v2-pro
Live JSON: `GET https://kymaapi.com/v1/models` (no auth required)

## Facts

- **Model ID**: `kling-avatar-v2-pro`: pass this as `model` in the request body
- **Creator**: Kuaishou
- **Released**: 2025-12-01
- **Price**: $0.15525 / second

## Call it

```bash
curl https://kymaapi.com/v1/videos/generations \
  -H "Authorization: Bearer $KYMA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "kling-avatar-v2-pro", "prompt": "..."}'
```

## Positioning

Kling AI Avatar v2 Pro is Kuaishou’s audio-driven avatar model. Give it one portrait and one speech or song track, and it returns a clip of that face lip-syncing to your audio.

## About

Kling AI Avatar v2 Pro, from Kuaishou’s Kling Avatar 2.0 line (late 2025), animates a single still image from an audio track. The mouth follows the speech or singing, the head and shoulders move naturally, and the audio you send plays through unchanged. There is no duration setting: the clip runs as long as the audio.

In our own tests it kept a synthetic singer’s identity, lip-synced a Vietnamese vocal with the mouth closing on breaths, and kept the original vocal in the soundtrack. A 13 second clip took about six minutes. Kyma accepts uncompressed PCM WAV audio of at least 5 seconds (WAV PCM only for now).

Billing is per second of the finished clip. You declare your audio’s length as duration_s; Kyma checks the audio before generating, cuts anything longer, and charges the clip’s real length.

## Use cases

- **Talking-Head Clips**, Turns a presenter photo and a voice recording into a speaking clip for explainers, lessons and announcements.
- **Lip-Synced Singing**, Makes a character sing your own song track, keeping the real vocal in the soundtrack.
- **Any-Language Voiceover**, Lip-syncs to the audio itself rather than to a transcript, so the language of the recording does not need to be supported separately.
- **Character Series**, Reuses one still of a character across many clips, each driven by a different line of audio.

## Not ideal for

Do not use this model for text-to-video scenes, camera moves, or clips longer than 60 seconds; it animates one face from one image and needs an audio track.

## See also

- All models: https://kymaapi.com/models.md
- Pricing: https://kymaapi.com/pricing.md
- Other models by Kuaishou: https://kymaapi.com/models?q=Kuaishou
