> ## Documentation Index
> Fetch the complete documentation index at: https://docs.comfy.org/llms.txt
> Use this file to discover all available pages before exploring further.

# WanImageToVideo - ComfyUI Built-in Node Documentation

> The WanImageToVideo node prepares conditioning and latent representations for video generation.

The WanImageToVideo node prepares conditioning and latent representations for video generation. It creates an empty latent space for the video and can optionally incorporate a starting image and CLIP vision output to guide the generation. Both the positive and negative conditioning inputs are updated with the provided image and vision data.

## Inputs

| Parameter | Description | Data Type | Required | Range |
| - | - | - | - | - |
| `positive` | Positive conditioning input used to guide the generation | CONDITIONING | Yes | - |
| `negative` | Negative conditioning input used to guide the generation | CONDITIONING | Yes | - |
| `vae` | VAE model used to encode images into latent space | VAE | Yes | - |
| `width` | Width of the generated video (default: 832, step: 16) | INT | Yes | 16 to MAX\_RESOLUTION |
| `height` | Height of the generated video (default: 480, step: 16) | INT | Yes | 16 to MAX\_RESOLUTION |
| `length` | Number of frames in the video (default: 81, step: 4) | INT | Yes | 1 to MAX\_RESOLUTION |
| `batch_size` | Number of videos to generate in one batch (default: 1) | INT | Yes | 1 to 4096 |
| `clip_vision_output` | Optional CLIP vision output added as extra conditioning to both the positive and negative inputs | CLIP\_VISION\_OUTPUT | No | - |
| `start_image` | Optional starting image used to initialize the video. When provided, it is resized to the specified `width` and `height` and placed at the beginning of the frame sequence; any frames beyond `length` are ignored. The remaining frames are filled with neutral gray (0.5) values unless `ref_pad_image` is supplied. | IMAGE | No | - |
| `ref_pad_image` | Optional reference image whose first frame replaces the neutral gray padding of the frame sequence. It is resized to the specified `width` and `height`, and only the first image of the batch is used. Has an effect only when `start_image` is also provided. | IMAGE | No | - |

**Note:** When `start_image` is provided, the frame sequence is encoded with the VAE and a mask is applied to the conditioning. The mask is set to 0 for the frames covered by the starting image and 1 for the remaining frames, so generation continues from the provided image. Only the first three color channels (RGB) of the image are used during encoding. Both positive and negative conditioning receive the same concatenated latent image, mask, and (if supplied) CLIP vision output. When `ref_pad_image` is supplied together with `start_image`, its first frame is resized to `width` and `height` and written into the RGB channels of the padding frames before the starting image is placed on top, so the padding carries the reference image instead of flat gray. This is SVI-style anti-drift padding, used by models such as ID-V2V.

## Outputs

| Output Name | Description | Data Type |
| - | - | - |
| `positive` | Positive conditioning, updated with the image and vision data | CONDITIONING |
| `negative` | Negative conditioning, updated with the image and vision data | CONDITIONING |
| `latent` | Empty latent tensor ready for video generation, with shape \[batch\_size, 16, ((length-1)//4)+1, height//8, width//8] | LATENT |

> This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! [Edit on GitHub](https://github.com/Comfy-Org/embedded-docs/blob/main/comfyui_embedded_docs/docs/WanImageToVideo/en.md)

***

**Source fingerprint (SHA-256):** `3000c1c816d2c123fc5bc46ea1f193c52c0f81a2f8c4a9110d8a4fa909185aea`


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.