# Wan 2.2 A14B Speech to Video Turbo API

> Wan 2.2 A14B Speech to Video Turbo — high-quality AI video generation.

- **Provider**: Wan
- **Model id**: `wan-2-2-a14b-speech-to-video-turbo`
- **Modality**: video
- **Price**: 8.21–16.42 credits /sec

## Overview

Wan 2.2 A14B Speech to Video Turbo is called in two steps: create a generation task, then poll the task until the result is ready.

## Authentication

All requests require a Bearer Token in the request header:

```
Authorization: Bearer YOUR_API_KEY
```

## Create Task

`POST https://you.bot/api/v1/generate`

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| modelId | string | Yes | Model id: `wan-2-2-a14b-speech-to-video-turbo` |
| input | object | Yes | Input parameters object (see below) |
| callbackUrl | string | No | https URL we POST the finished task to. Signed with `X-Webhook-Signature` once you create a webhook signing key in Dashboard → Settings |

### input object parameters

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| prompt | string | Yes | The text prompt used for video generation Max 5000 characters. (example: Subtle lifelike motion driven by the audio, gentle natural head movement, soft cinematic lighting, smooth realistic delivery.) |
| image_url | string | Yes | URL of the input image. If the input image does not match the chosen aspect ratio, it is resized and center cropped (image URL) |
| audio_url | string | Yes | The URL of the audio file (image URL) |
| num_frames | number | No | Number of frames to generate. Must be between 40 to 120, (must be multiple of 4) (range 40-120) (default: 80) |
| frames_per_second | number | No | Frames per second of the generated video. Must be between 4 to 60. When using interpolation and adjust_fps_for_interpolation is set to true (default true,) the final FPS will be multiplied by the number of interpolated frames plus one. For example, if the generated frames per second is 16 and the number of interpolated frames is 1, the final frames per second will be 32. If adjust_fps_for_interpolation is set to false, this value will be used as-is (range 4-60) (default: 16) |
| resolution | string | No | Resolution of the generated video (480p, 580p, or 720p) (options: 480p \| 580p \| 720p) (default: 480p) |
| negative_prompt | string | No | Negative prompt for video generation Max 500 characters. |
| seed | number | No | Random seed for reproducibility. If None, a random seed is chosen |
| num_inference_steps | number | No | Number of inference steps for sampling. Higher values give better quality but take longer (range 2-40) (default: 27) |
| guidance_scale | number | No | Classifier-free guidance scale. Higher values give better adherence to the prompt but may decrease quality (range 1-10) (default: 3.5) |
| shift | number | No | Shift value for the video. Must be between 1.0 and 10.0 (range 1-10) (default: 5) |
| enable_safety_checker | boolean | No | If set to true, input data will be checked for safety before processing (true/false) (default: true) |
| nsfw_checker | boolean | No | A configurable parameter. Defaults to true in the Playground. (true/false) (default: true) |

### Request example

```json
{
  "modelId": "wan-2-2-a14b-speech-to-video-turbo",
  "input": {
    "prompt": "Subtle lifelike motion driven by the audio, gentle natural head movement, soft cinematic lighting, smooth realistic delivery.",
    "image_url": "https://you.bot/examples/grok-imagine-text-to-image.jpg",
    "audio_url": "https://you.bot/examples/suno-v4-5.mp3",
    "num_frames": 80,
    "frames_per_second": 16,
    "resolution": "480p",
    "negative_prompt": "blurry, low quality, distorted, watermark, text",
    "seed": 1,
    "num_inference_steps": 27,
    "guidance_scale": 3.5,
    "shift": 5,
    "enable_safety_checker": true,
    "nsfw_checker": true
  }
}
```

### Response example

```json
{
  "taskId": "281e5b0…f39b9",
  "creditsCharged": 41.05
}
```

## Query Task

`GET https://you.bot/api/v1/task/{taskId}?model=wan-2-2-a14b-speech-to-video-turbo`

When `state` is `success`, the output is in `resultUrls`: `{ "state": "success", "resultUrls": [ ... ] }`. (Text models return inline in the create response.)

## Error Codes

| Code | Description |
|------|-------------|
| 200 | Request successful |
| 400 | Invalid request parameters |
| 401 | Authentication failed — check API Key |
| 402 | Insufficient account balance |
| 403 | IP not allowed for this key, or the key is restricted to other models |
| 404 | Task not found, or it has expired |
| 409 | Duplicate request — an identical submit arrived moments ago. Nothing was charged; retry is safe |
| 413 | Payload too large — see the character and file-size limits for this model |
| 429 | Rate limit exceeded — retry after the Retry-After header |
| 500 | Internal server error |
| 504 | Upstream timed out — nothing was delivered; retry is safe |
