Skip to main content
MiniMax H3 Reference to Video creates the text conditioning and the empty audio-video latent needed for MiniMax H3 reference-to-video generation. You provide a prompt plus optional reference images, videos, and audio clips, and the node encodes these references into conditioning the model can use while generating. The prompt refers to the references with <Picture i>, <Video k>, and <Audio j> tags.

Inputs

Notes:
  • The prompt refers to reference media with 1-based tags per type: <Picture i> for images, <Video k> for videos, and <Audio j> for audio. References are presented to the model in a fixed order: images, then videos (with each soundtrack’s <Audio j> label right before its <Video k>), then standalone audio.
  • A soundtrack connected to ref_video_audio_N is used with the reference video connected to ref_video_N.
  • Reference videos must contain at least 5 frames (~0.2 seconds at 24 fps), otherwise the node raises an error. Frames beyond the requested length are trimmed, and the remaining frame count is adjusted to a value supported by the model.
  • The requested length is aligned to a supported frame count before the latent is created.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 47df0d6d13cb02aa4f69b50a7f8d0f6c1639c1fb5e0f69bf8fc57dd4cb752db8