Inputs
Note: When
start_image is provided, the frame sequence is encoded with the VAE and a mask is applied to the conditioning. The mask is set to 0 for the frames covered by the starting image and 1 for the remaining frames, so generation continues from the provided image. Only the first three color channels (RGB) of the image are used during encoding. Both positive and negative conditioning receive the same concatenated latent image, mask, and (if supplied) CLIP vision output. When ref_pad_image is supplied together with start_image, its first frame is resized to width and height and written into the RGB channels of the padding frames before the starting image is placed on top, so the padding carries the reference image instead of flat gray. This is SVI-style anti-drift padding, used by models such as ID-V2V.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
3000c1c816d2c123fc5bc46ea1f193c52c0f81a2f8c4a9110d8a4fa909185aea