Skip to main content
API Reference for Wan 3.0 Video. Wan 3.0 Video generates from a prompt alone or from a media list of first and last frames, reference images, videos and audio, addressed in the prompt as Image 1, Video 1 and Audio 1. Input: Text, Image, Video, Audio · Output: Video + Audio

Quick start

Create a key in your Comfy workspace and export it as COMFY_API_KEY. The Python and TypeScript snippets use the Comfy SDKs (pip install comfy-sdk and npm install @comfyorg/sdk); the cURL snippet is the same call over raw HTTP. Model ID: wan/wan3.0-video Endpoint: POST https://api.comfy.org/v2/models/wan/wan3.0-video
The same body, sent to POST https://api.comfy.org/v2/models/wan/wan3.0-video/requests. Router answers 201 with a request_id as soon as the run is admitted, and the result is collected once it is ready, from this process or another one. Queued delivery walks through status, cancellation and collection.

Serving providers

This model is served by Comfy Router directly unless the request names another provider. The providers below serve it too, on the same endpoint and with the same model ID, selected with the model_provider query parameter.
  • Comfy (default): POST https://api.comfy.org/v2/models/wan/wan3.0-video
  • Higgsfield, as higgsfield/higgsfield-wan-3: POST https://api.comfy.org/v2/models/wan/wan3.0-video?model_provider=higgsfield
strict_mode defaults to false, so Router translates the native request body documented on this page into the provider’s own schema and translates the response back. See model_provider, strict_mode and fallback_provider in the API reference, and Serving providers for every model routed this way.

Schema

Input

object
required
Enter basic information, such as prompt words, etc.
string
Audio file download URL. Supported formats: mp3 and wav. Cannot be used with reference_video_urls.
string
First frame image URL or Base64 encoded data. Required for the wan2.5-i2v-preview and wan2.6-i2v models. The happyhorse-1.x i2v spellings do NOT take their first frame here — they take it as a media element of type first_frame (see media below), and an img_url body that succeeds on wan2.6-i2v is refused by the provider on happyhorse-1.0-i2v and happyhorse-1.1-i2v — and the wan2.7/wan3.0 spellings take their reference assets in media as well. Image formats: JPEG, JPG, PNG, BMP, WEBP. Resolution: 360-2000 pixels. File size: max 10MB.
object[]
Media asset list. Specifies reference materials (image, audio, video) for video generation. Each element contains a type and url field. This field is shared by every model served from this component, and which type values a given model accepts — and whether it carries its image or video input here at all rather than in img_url or reference_video_urls — varies by model. Supported type values, per model:
  • wan2.7-i2v: first_frame, last_frame, driving_audio, first_clip. first_frame and last_frame: JPEG/JPG/PNG/BMP/WEBP, 240-8000 pixels per side, aspect ratio between 1:8 and 8:1, at most 20MB, a public URL or a data:{MIME_type};base64,… URL. driving_audio (WAV/MP3, 2-30 seconds, at most 15MB) and first_clip (MP4/MOV, 2-10 seconds, 240-4096 pixels per side, at most 100MB): a public HTTP/HTTPS URL; an inline data: URL is not accepted.
  • wan2.7-r2v: first_frame (optional, max 1), reference_image, reference_video. At least one reference_image or reference_video is required, and reference images plus reference videos total at most 5. first_frame and reference_image: JPEG/JPG/PNG/BMP/WEBP (no alpha channel), 240-8000 pixels per side, aspect ratio between 1:8 and 8:1, at most 20MB, a public URL or a data:{MIME_type};base64,… URL. reference_video: a public HTTP/HTTPS URL; an inline data: URL is not accepted.
  • wan2.7-videoedit: video (exactly 1, the clip being edited: MP4/MOV, 2-10 seconds, 240-4096 pixels per side, at most 100MB, a public HTTP/HTTPS URL; an inline data: URL is not accepted) plus an optional 0 to 4 reference_image (JPEG/JPG/PNG/BMP/WEBP, 240-8000 pixels per side, aspect ratio between 1:8 and 8:1, at most 20MB, a public URL or a data:{MIME_type};base64,… URL).
  • wan3.0-video, wan3.0-video-prime: first_frame (max 1), last_frame (max 1), reference_image (max 10), reference_video (max 5 clips, total duration <= 15s), reference_audio (max 5 clips, total duration <= 15s), file (max 1, cannot be used with link), link (max 1, cannot be used with file). The reference_*/file/link types and first_frame/last_frame types are mutually exclusive within the same request. The array order defines the reference order of assets in the prompt (Image 1, Video 1, Audio 1, …). first_frame, last_frame and reference_image: JPEG/JPG/PNG/BMP/WEBP, 240-8000 pixels per side, aspect ratio at most 8:1, at most 20MB, a public URL or a data:{MIME_type};base64,… URL. reference_video, reference_audio, file and link: a public HTTP/HTTPS URL; an inline data: URL is not accepted.
  • happyhorse-1.0-i2v, happyhorse-1.1-i2v: first_frame only, exactly 1. These spellings take their first frame HERE and not in img_url. At least 300x300 pixels, aspect ratio between 1:2.5 and 2.5:1, JPEG/JPG/PNG/WEBP, at most 20MB, a public URL or a data:{MIME_type};base64,… URL.
  • happyhorse-1.0-r2v, happyhorse-1.1-r2v: reference_image only, 1 to 9 of them. This is the model’s only reference input. Shortest side at least 400 pixels, JPEG/JPG/PNG/WEBP, at most 20MB, a public URL or a data:{MIME_type};base64,… URL. reference_video is not an input type for this operation.
  • happyhorse-1.0-video-edit: video (exactly 1, the clip being edited: MP4/MOV, 3-60 seconds, longer side at most 4096 pixels, shorter side at least 360 pixels, aspect ratio between 1:2.5 and 2.5:1, at most 100MB, a public HTTP/HTTPS URL; an inline data: URL is not accepted) plus an optional 0 to 5 reference_image (JPEG/JPG/PNG/WEBP, at least 300x300 pixels, aspect ratio between 1:2.5 and 2.5:1, at most 20MB, a public URL or a data:{MIME_type};base64,… URL).
  • happyhorse-1.0-t2v, happyhorse-1.1-t2v, wan2.7-t2v, wan2.6-t2v, wan2.5-t2v-preview: text-to-video operations. Their partner references document no media input, so leave this field unset.
  • wan2.5-i2v-preview, wan2.6-i2v, wan2.6-r2v: this component enumerates no media vocabulary for these spellings. wan2.5-i2v-preview and wan2.6-i2v take their first frame in img_url; wan2.6-r2v takes its reference clips in reference_video_urls. The per-model counts above are the partner’s constraints, not this schema’s: no minItems/maxItems is declared on media, because the admissible types and their caps differ per model. Schema acceptance does not imply provider acceptance — and, in the other direction, the per-asset “at most 20MB” figures above — the ceiling on the image inputs, the only inputs that may be sent inline — are the partner’s ceiling on the asset it ends up with, and they apply as written when the asset is a public URL, because those bytes never travel through Comfy. An inline data:{MIME_type};base64,… URL does travel through Comfy and meets a transport ceiling as well: a Comfy Router POST body is capped at 100 MiB in total and answered 413 past it, and base64 inflates a payload by about 4/3, so a single inline asset above roughly 75 MB is refused before any of the partner rules here are reached. At that size a single asset at the partner’s own 20MB ceiling fits inline comfortably, so the partner rule is what binds for one asset — the transport ceiling binds across several of them, since the 100 MiB Router cap bounds the WHOLE request rather than each element. Send anything near these ceilings as a public URL.
string
required
Media asset typePossible values: first_frame, last_frame, driving_audio, first_clip, reference_image, reference_video, reference_audio, video, file, link
string
required
URL of the media file: a public HTTP/HTTPS URL, or — for exactly the IMAGE inputs the media description above lists for each model: first_frame and last_frame on wan2.7-i2v; first_frame and reference_image on wan2.7-r2v; reference_image on wan2.7-videoedit; first_frame, last_frame and reference_image on wan3.0-video and wan3.0-video-prime; first_frame on happyhorse-1.0-i2v and happyhorse-1.1-i2v; reference_image on happyhorse-1.0-r2v and happyhorse-1.1-r2v; and reference_image on happyhorse-1.0-video-edit — an inline data:{MIME_type};base64,... URL. No video, audio, file or link input on any of these models takes an inline data: URL. An oss://dashscope-instant temporary URL from the partner’s own upload service, or (on wan2.7-videoedit) an Asset Center asset_id, resolves only against the Alibaba Cloud account that uploaded it; Comfy makes the model call with its own key and exposes no upload route, so a caller’s own oss:// URL or asset_id is not a supported input. This field does not validate the URL scheme, so such a value is forwarded to the partner as-is rather than rejected here. See the media description for the per-model size and pixel limits and for the 100 MiB Router request-body cap that bounds an inline payload in aggregate.
string
Reverse prompt words are used to describe content that you do not want to see in the video screen
string
Text prompt words. Support Chinese and English, length not exceeding 800 characters (up to 20,000 characters for wan3.0-video; content exceeding the limit is truncated). For wan2.6-r2v with multiple reference videos, use ‘character1’, ‘character2’, etc. to refer to subjects in the order of reference videos. Example: “Character1 sings on the roadside, Character2 dances beside it” For wan3.0-video reference mode, use ‘Image 1’, ‘Video 1’, ‘Audio 1’, etc. to refer to media assets in the corresponding order within the media array.
string[]
Reference video URLs for wan2.6-r2v model only. Array of 1-3 video URLs. Input restrictions:
  • Format: mp4, mov
  • Quantity: 1-3 videos
  • Single video length: 2-30 seconds
  • Single file size: max 30MB
  • Cannot be used with audio_url Reference duration: Single video max 5s, two videos max 2.5s each, three videos proportionally less. Billing: Based on actual reference duration used.
string
Video effect template name. Optional. Currently supported: squish, flying, carousel. When used, prompt parameter is ignored.
string
The ID of the model to call. NOT constrained on this component: Comfy Router fills it from the {model} path segment of POST /v2/models/wan/{model}, so a Router caller omits it. A direct v1 call to POST /proxy/wan/api/v1/services/aigc/video-generation/video-synthesis MUST supply it, and the enum of accepted spellings lives on that operation’s own component, WanVideoGenerationRequest.
object
Video processing parameters
boolean
default:"true"
Whether to add audio to the video. Supported for wan3.0-video and wan3.0-video-prime only; the provider documents no such switch for the other wan models.
string
default:"\"auto\""
Video audio setting for wan2.7-videoedit model.
  • auto (default): Model intelligently judges based on prompt content
  • origin: Forcefully preserve the original audio from the input video Possible values: auto, origin
integer
default:"5"
The duration of the video generated, in seconds:
  • wan2.5 models: 5 or 10 seconds
  • wan2.6-t2v, wan2.6-i2v: 5, 10, or 15 seconds
  • wan2.6-r2v: 5 or 10 seconds only (no 15s support)
  • wan2.7-i2v, wan2.7-t2v: integer in [2, 15]
  • wan2.7-r2v, wan2.7-videoedit: integer in [2, 10]
  • wan3.0-video: integer in [2, 30] without video input; with video input the total input video duration + output video duration must not exceed 30 seconds; -1 enables smart duration mode where the model picks a suitable duration This proxy returns 400 for a duration outside [-1, 30]. Range: -1 to 30
boolean
default:"true"
Is it enabled prompt intelligent rewriting. Default is true
string
Aspect ratio of the generated video. For wan2.7 and wan3.0 models only. For wan2.7 models, defaults based on the resolution tier if not provided. For wan3.0-video, adaptive (the default) automatically recommends a suitable aspect ratio based on the input media proportions and intent.Possible values: adaptive, 16:9, 9:16, 1:1, 4:3, 3:4
string
Resolution level. Supported values vary by model:
  • wan2.5-i2v-preview: 480P, 720P, 1080P
  • wan2.6 models (t2v, i2v, r2v): 720P, 1080P only (no 480P support)
  • wan2.7 models (i2v, t2v, r2v, videoedit): 720P, 1080P (default 1080P)
  • wan3.0-video, wan3.0-video-prime: 480P, 720P, 1080P (upstream default 1080P)
  • happyhorse-1.0-* (t2v, i2v, r2v, video-edit): 720P, 1080P
  • happyhorse-1.1-* (t2v, i2v, r2v): 480P, 720P, 1080P This proxy rejects video generation requests that provide neither resolution nor size, because the resolution tier selects the billing rate. It upper-cases resolution before checking it and forwards the upper-cased value, returns 400 for a value outside this enum, and returns 400 for 480P (whether sent as resolution or as a 480P size) on the wan2.6, wan2.7 and happyhorse-1.0 models listed above as having no 480P. Possible values: 480P, 720P, 1080P
integer
Random number seed, used to control the randomness of the model generated contentRange: 0 to 2147483647
string
default:"\"single\""
Intelligent multi-lens control. Only active when prompt_extend is enabled. For wan2.6 and wan2.7-r2v models.
  • single: Single-shot video (default)
  • multi: Multi-shot video Possible values: multi, single
string
Video resolution in format widthheight. Supported resolutions vary by model: For wan2.5 T2V: 480P (480832, 832480, 624624), 720P, 1080P sizes For wan2.6 T2V/R2V (no 480P): 720P: 1280720, 7201280, 960960, 1088832, 8321088 1080P: 19201080, 10801920, 14401440, 16321248, 12481632 When both size and resolution are sent, size selects the billed tier, and resolution must still be a value this proxy accepts for the model. A 480P size returns 400 on every model the resolution description lists as having no 480P.
boolean
default:"false"
Whether to add a watermark logo, the watermark is located in the lower right corner
Generated from the schema Router serves at GET /v2/models/wan/wan3.0-video/openapi.json, the same document it validates a call against before the request reaches the provider.

Output

object
required
string
Actual prompt after intelligent rewriting (for video tasks)
string
Audio URL for I2V tasks with audio generation
string
The error code for the failed request (not returned if request is successful)
string
Task completion time
string
Detailed information about the failed request (not returned if request is successful)
string
Original input prompt (for video tasks)
object[]
List of task results for image generation tasks
string
Actual prompt after intelligent rewriting (if enabled)
string
Image error code (returned when some tasks fail)
string
Image error information (returned when some tasks fail)
string
Original input prompt
string
Generated image URL address
string
Task execution time
string
Task submission time
string
required
Task ID
object
Task result statistics for image generation tasks
integer
Number of failed tasks
integer
Number of successful tasks
integer
Total number of tasks
string
required
Task statusPossible values: PENDING, RUNNING, SUCCEEDED, FAILED, CANCELED, UNKNOWN
string
Video URL for completed video generation tasks. Link validity period 24 hours
string
required
Unique request identifier
object
Output information statistics. Only successful results are counted
integer
Video resolution level (I2V and wan3.0-video tasks)
number
Duration of generated video in seconds (I2V and wan3.0-video tasks)
integer
Frame rate of the generated video (wan3.0-video tasks)
integer
Number of generated images (T2I and I2I tasks)
number
Duration of the input video in seconds, 0.0 when no video input (wan3.0-video tasks)
number
Duration of the output video in seconds (wan3.0-video tasks)
string
Aspect ratio of the generated video, e.g. 16:9 (wan3.0-video tasks)
string
Image resolution (T2I and I2I tasks)
integer
Number of generated videos (T2V tasks)
number
Duration of generated video in seconds (T2V tasks)
string
Video resolution ratio (T2V tasks)
string
Error code for a failed request, reported at the ROOT of the envelope rather than under output (not returned if the request succeeded).
string
Detailed information about a failed request, reported at the ROOT of the envelope rather than under output (not returned if the request succeeded). Read this before falling back to output.message.

Examples

Input

Output

The video URL is valid for 24 hours. Download the video promptly if you need to keep it.

Before you ship

The SDKs create an Idempotency-Key and reuse it for automatic retries. For manual retries, reuse the original key. Router can hold the connection for up to 10 minutes. When a request fails, Router sends an X-Comfy-Error-Type response header explaining why. A 422 means Router rejected the input before calling the provider, and a 413 means the request body was larger than Router accepts. Download generated assets promptly because result URLs can expire. Any size limit named in a field description above is the provider’s own bound on that field, quoted from the provider’s specification. Router applies a separate cap to the whole request body, which base64-encoded media counts against: see request body size. This page documents one partner model called through Comfy Router. The same comfy-sdk / @comfyorg/sdk package also ships a second client, for running a whole ComfyUI workflow graph on Comfy Cloud: Comfy(api_key=...) / new Comfy({ apiKey }), with client.workflows, client.assets and client.jobs. See Comfy SDKs.

Headers

Authentication, idempotency, request IDs, error buckets, retry pacing, spend limits.

Using the Router API

Model discovery, validation errors, retries, and billing.

Limitations

What Router does not do today, and what to use instead.