> ## Documentation Index
> Fetch the complete documentation index at: https://docs.comfy.org/llms.txt
> Use this file to discover all available pages before exploring further.

# ComfyUI MiniMax H3 Prompt Guide: Official Guides, Tips, and Prompt Embeddings

> Write better MiniMax H3 prompts in ComfyUI: MiniMax's official prompt guides, general prompting tips, dialogue and speaker rules, and style embeddings.

The prompt is where H3's multimodal training pays off: shots, camera moves, dialogue, and sound effects all live in one prompt block. This page collects the prompt writing resources for MiniMax H3 in ComfyUI: MiniMax's official guides, general tips that apply to every workflow, the rules for dialogue and speakers, and prompt embeddings for style effects.

## Official guides

MiniMax publishes two official prompt writing guides:

* [Base generation modes](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md) (T2VA, I2VA, FL2VA, L2VA): structure a prompt into timed shots with camera movement and audio (dialogue, SFX, music), with examples for each mode
* [Full-reference mode](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md) (R2V): the rewrite output structure, including subject definitions, reference labels, and retention analysis, and how to assign each reference a role in the target shot

MiniMax also publishes installable [H3 skills](https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills), including a prompt-writing skill that packages these guides for agent use.

The workflow pages show these guides applied to each template: [T2V and I2V](/tutorials/video/minimax/minimax-h3-native), [R2V](/tutorials/video/minimax/minimax-h3-native#minimax-h3-reference-to-video-r2v), [Multiframe Reference](/tutorials/video/minimax/minimax-h3-multiframe#prompt-writing-guide), and [Fun ControlNet Union](/tutorials/video/minimax/minimax-h3-fun-controlnet#prompting-tips).

## General tips

The final prompt follows a fixed structure. I2VA prompts open with a fixed image-alignment sentence on the first line, followed by one blank line, for example `For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.` FL2VA and L2VA use their own fixed sentences in that line, which align each reference picture to a mark in the target video, and T2VA prompts have no instruction line and start directly with the fields. The body carries three fields in this order:

```text theme={null}
integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...
```

Write the prompt itself in English. Keep dialogue and lyrics inside `<d>` tags, and screen-visible text in its original language, in both cases verbatim.

1. **Describe the whole scene**: State the overall scene first (location, character, what is happening), then break it into timed shots. The overall style and the opening composition belong at the start of `[Shot 1]` inside `integrated_multimodal_description`, after any image-alignment instruction.
2. **Shots, camera, and audio**: Describe the shots, camera moves, and the accompanying audio (dialogue, SFX, music) in one prompt block
3. **Resolution**: H3's native canvas is a 768px short edge, which is 1344x768 at 16:9, and resolutions are rounded to a multiple of 32. See [Setting the output resolution](/tutorials/video/minimax/minimax-h3#setting-the-output-resolution)
4. **Length**: The node's `length` input is a frame count at 24fps (default 124, about 5 seconds) that snaps up to the model's 17-frame block grid (17k+5)
5. **On-screen text in quotes**: Wrap any text that is visible on screen (signs, banners, subtitles, neon text) in English double quotation marks, and list each string separately. Copy the original text and punctuation verbatim, without translating it, for example `A red neon sign reading "营业中" glows above the doorway.` Keep spoken dialogue in `<d>` tags instead: quotes are for printed text, `<d>` is for what characters say
6. **Audio reuse markers in R2V**: In full-reference (R2V) prompts, declare each audio reference's relationship in the retention analysis: `fully_copy` uses the source audio 1:1 as the final audio track, `partially_copy` copies part of the timeline or selected layers, or adds, removes, and replaces sounds afterwards, `reference` carries timbre, rhythm, music style, dialogue content, or sound texture without copying the signal, and `weak_reference` keeps only a broad similarity in category or atmosphere. The marker must match what the audio actually does in the target video. See the official full-reference guide's relationship marker table
7. **Write the audio lanes explicitly**: Each prompt carries two audio fields, and they follow different rules. `overall_soundscape` summarizes ambient sound, physical action sounds, and non-verbal human sounds in 1 to 4 sentences (dialogue, singing, and diegetic music belong in `integrated_multimodal_description`, not here), and uses `N/A` only when the clip should be completely silent. `non_diegetic_music` describes score that only the audience hears in 1 to 3 sentences, and uses `N/A` when there is no such music. If a quiet shot comes back with speech you did not ask for, write out both fields explicitly and regenerate. See the official base guide's sections on `overall_soundscape` and `non_diegetic_music`
8. **Write bans as positives**: the H3 templates sample through the `BasicGuider` node, which carries a single conditioning input, so there is no negative branch and a negative prompt has no effect (the guider runs at `cfg` 1, where ComfyUI skips the unconditional pass; a `CFGGuider` only uses negative conditioning above `cfg` 1). An instruction that names an unwanted element adds that wording to the description the model reads, with no separate pass to subtract it, so a line such as `no subtitles and no on-screen text` registers text as content. Say what should be in the shot instead, for example `the sign above the door is blank`. In full-reference prompts, scope the same thing through the retention analysis, where `fully_preserved`, `partially_preserved`, `attribute_transfer`, and `weak_reference` state what each reference contributes to the target video
9. **Reference sheets (R2V)**: one reference image can hold several views of a subject or the panels of a storyboard, and the prompt then addresses one panel at a time. Give the sheet its own `<Picture N>` entry rather than citing it inside a `<Subject N>` definition, since a picture cited only inside a definition is not used as a separate reference, and describe the sheet there, for example `<Picture 2> is a character sheet with three panels: a portrait, a full front view, and a full back view of the woman`. Point a shot at one panel inside the shot description, for example `[Shot 2] the shot's keyframe corresponds to the full front view panel of <Picture 2>`. A sheet keeps several views inside the local R2V node's limit of 9 reference images and lets each shot name the view it needs. Number or label the panels so the prompt can point at them, and keep the sheet's printed text to those labels.

Workflow-specific prompting (reference tags, timeline anchors, control-video length) is covered on each workflow page.

## Dialogue and speakers

Spoken lines follow fixed rules in every workflow, and they come from the official base guide's section on speakers, dialogue, and singing.

* **Stable speaker IDs**: give every voice in the prompt an ID such as `(S1)` or `(S2)`, including off-screen and singing voices, and repeat that ID in every shot where it speaks. The IDs number speakers, not subjects: a character who never vocalizes gets no ID, and the first voice in the prompt is `(S1)` even when its character is numbered differently. Speakers who vocalize together take a compound ID such as `(S1,S2)`
* **Outside and inside `<d>`**: the identifying phrase, the speaker ID, the action, and the delivery stay outside the tags, and the tags carry only the language tag plus the words themselves, copied verbatim. Write each line inside the shot where it is spoken, in the same paragraph as that shot's action
* **Voiceover**: a line the audience hears without the character speaking it needs the exact phrase `says in an off-screen voiceover`, plus a statement right after the `<d>` block that the on-screen character's lips remain closed. A line that continues across a cut needs `<scenetrans>` at both connection points and a statement that the audio carries over the transition
* **Several speakers with audio references**: in full-reference (R2V) prompts that pair two or more speakers with audio references, H3 can hand an audio reference to the wrong speaker even when the prompt text and the reference connection order are both correct. Keep each shot to the references that shot needs, and tie a line to a visible event, for example `When the phone is at his ear, the man in the coat speaks`, rather than to an absolute timecode, which is what cut times are for. When a line still lands on the wrong speaker, generate it with a voice tool and supply it as that speaker's audio reference, with the relationship marker that matches the clip: `fully_copy` when the clip becomes the video's complete final audio track, or `partially_copy` when it covers part of the timeline

## Prompt embeddings

[Comfy-Org/ComfyUI#15697](https://github.com/Comfy-Org/ComfyUI/pull/15697) added support for prompt embeddings in MiniMax H3. You can use ComfyUI's standard `embedding:` syntax in H3 prompts.

Place an embedding file in `ComfyUI/models/embeddings/` and reference it in the prompt by name, for example `embedding:my_embedding`. The embedding is loaded and mixed into the text conditioning just like with any other ComfyUI model.

The [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/embeddings) repository hosts 10 style embeddings in its `embeddings` folder. These files are unofficial: they were contributed by community member [silveroxides](https://huggingface.co/silveroxides) via [Hugging Face PR #50](https://huggingface.co/Comfy-Org/MiniMax-H3/discussions/50), and were not produced by Comfy-Org or MiniMax. The original files are in the [silveroxides/MiniMax-H3\_tests](https://huggingface.co/silveroxides/MiniMax-H3_tests/tree/main/embeddings) repository. The file name describes the intended effect, and the trigger word is the file name without the extension:

| Embedding | Effect (by name) |
| - | - |
| `minimaxh3_art_is_explosion` | Explosive art composition |
| `minimaxh3_blooming_flowers` | Flowers blooming |
| `minimaxh3_bullet_time` | Bullet-time effect |
| `minimaxh3_dark_magic` | Dark magic atmosphere |
| `minimaxh3_fire_breath` | Fire-breath effect |
| `minimaxh3_four_seasons` | Four seasons transition |
| `minimaxh3_kiss_camera` | Kiss scene camera move |
| `minimaxh3_spiral_ascent` | Spiraling ascent |
| `minimaxh3_storm_magic` | Storm magic |
| `minimaxh3_truman_show` | Truman Show style |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.