Official guides
MiniMax publishes two official prompt writing guides:- Base generation modes (T2VA, I2VA, FL2VA, L2VA): structure a prompt into timed shots with camera movement and audio (dialogue, SFX, music), with examples for each mode
- Full-reference mode (R2V): the rewrite output structure, including subject definitions, reference labels, and retention analysis, and how to assign each reference a role in the target shot
General tips
The final prompt follows a fixed structure. I2VA prompts open with a fixed image-alignment sentence on the first line, followed by one blank line, for exampleFor the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. FL2VA and L2VA use their own fixed sentences in that line, which align each reference picture to a mark in the target video, and T2VA prompts have no instruction line and start directly with the fields. The body carries three fields in this order:
<d> tags, and screen-visible text in its original language, in both cases verbatim.
- Describe the whole scene: State the overall scene first (location, character, what is happening), then break it into timed shots. The overall style and the opening composition belong at the start of
[Shot 1]insideintegrated_multimodal_description, after any image-alignment instruction. - Shots, camera, and audio: Describe the shots, camera moves, and the accompanying audio (dialogue, SFX, music) in one prompt block
- Resolution: H3’s native canvas is a 768px short edge, which is 1344x768 at 16:9, and resolutions are rounded to a multiple of 32. See Setting the output resolution
- Length: The node’s
lengthinput is a frame count at 24fps (default 124, about 5 seconds) that snaps up to the model’s 17-frame block grid (17k+5) - On-screen text in quotes: Wrap any text that is visible on screen (signs, banners, subtitles, neon text) in English double quotation marks, and list each string separately. Copy the original text and punctuation verbatim, without translating it, for example
A red neon sign reading "营业中" glows above the doorway.Keep spoken dialogue in<d>tags instead: quotes are for printed text,<d>is for what characters say - Audio reuse markers in R2V: In full-reference (R2V) prompts, declare each audio reference’s relationship in the retention analysis:
fully_copyuses the source audio 1:1 as the final audio track,partially_copycopies part of the timeline or selected layers, or adds, removes, and replaces sounds afterwards,referencecarries timbre, rhythm, music style, dialogue content, or sound texture without copying the signal, andweak_referencekeeps only a broad similarity in category or atmosphere. The marker must match what the audio actually does in the target video. See the official full-reference guide’s relationship marker table - Write the audio lanes explicitly: Each prompt carries two audio fields, and they follow different rules.
overall_soundscapesummarizes ambient sound, physical action sounds, and non-verbal human sounds in 1 to 4 sentences (dialogue, singing, and diegetic music belong inintegrated_multimodal_description, not here), and usesN/Aonly when the clip should be completely silent.non_diegetic_musicdescribes score that only the audience hears in 1 to 3 sentences, and usesN/Awhen there is no such music. If a quiet shot comes back with speech you did not ask for, write out both fields explicitly and regenerate. See the official base guide’s sections onoverall_soundscapeandnon_diegetic_music - Write bans as positives: the H3 templates sample through the
BasicGuidernode, which carries a single conditioning input, so there is no negative branch and a negative prompt has no effect (the guider runs atcfg1, where ComfyUI skips the unconditional pass; aCFGGuideronly uses negative conditioning abovecfg1). An instruction that names an unwanted element adds that wording to the description the model reads, with no separate pass to subtract it, so a line such asno subtitles and no on-screen textregisters text as content. Say what should be in the shot instead, for examplethe sign above the door is blank. In full-reference prompts, scope the same thing through the retention analysis, wherefully_preserved,partially_preserved,attribute_transfer, andweak_referencestate what each reference contributes to the target video - Reference sheets (R2V): one reference image can hold several views of a subject or the panels of a storyboard, and the prompt then addresses one panel at a time. Give the sheet its own
<Picture N>entry rather than citing it inside a<Subject N>definition, since a picture cited only inside a definition is not used as a separate reference, and describe the sheet there, for example<Picture 2> is a character sheet with three panels: a portrait, a full front view, and a full back view of the woman. Point a shot at one panel inside the shot description, for example[Shot 2] the shot's keyframe corresponds to the full front view panel of <Picture 2>. A sheet keeps several views inside the local R2V node’s limit of 9 reference images and lets each shot name the view it needs. Number or label the panels so the prompt can point at them, and keep the sheet’s printed text to those labels.
Dialogue and speakers
Spoken lines follow fixed rules in every workflow, and they come from the official base guide’s section on speakers, dialogue, and singing.- Stable speaker IDs: give every voice in the prompt an ID such as
(S1)or(S2), including off-screen and singing voices, and repeat that ID in every shot where it speaks. The IDs number speakers, not subjects: a character who never vocalizes gets no ID, and the first voice in the prompt is(S1)even when its character is numbered differently. Speakers who vocalize together take a compound ID such as(S1,S2) - Outside and inside
<d>: the identifying phrase, the speaker ID, the action, and the delivery stay outside the tags, and the tags carry only the language tag plus the words themselves, copied verbatim. Write each line inside the shot where it is spoken, in the same paragraph as that shot’s action - Voiceover: a line the audience hears without the character speaking it needs the exact phrase
says in an off-screen voiceover, plus a statement right after the<d>block that the on-screen character’s lips remain closed. A line that continues across a cut needs<scenetrans>at both connection points and a statement that the audio carries over the transition - Several speakers with audio references: in full-reference (R2V) prompts that pair two or more speakers with audio references, H3 can hand an audio reference to the wrong speaker even when the prompt text and the reference connection order are both correct. Keep each shot to the references that shot needs, and tie a line to a visible event, for example
When the phone is at his ear, the man in the coat speaks, rather than to an absolute timecode, which is what cut times are for. When a line still lands on the wrong speaker, generate it with a voice tool and supply it as that speaker’s audio reference, with the relationship marker that matches the clip:fully_copywhen the clip becomes the video’s complete final audio track, orpartially_copywhen it covers part of the timeline
Prompt embeddings
Comfy-Org/ComfyUI#15697 added support for prompt embeddings in MiniMax H3. You can use ComfyUI’s standardembedding: syntax in H3 prompts.
Place an embedding file in ComfyUI/models/embeddings/ and reference it in the prompt by name, for example embedding:my_embedding. The embedding is loaded and mixed into the text conditioning just like with any other ComfyUI model.
The Comfy-Org/MiniMax-H3 repository hosts 10 style embeddings in its embeddings folder. These files are unofficial: they were contributed by community member silveroxides via Hugging Face PR #50, and were not produced by Comfy-Org or MiniMax. The original files are in the silveroxides/MiniMax-H3_tests repository. The file name describes the intended effect, and the trigger word is the file name without the extension: