MiniMax H3 作為全模態影音生成模型,其提示詞撰寫邏輯與傳統純文生圖或一般影片模型有著顯著差異。在 H3 的架構中,提示詞必須同時承載時間軸發展、視覺鏡頭運算以及原生音效與背景音樂的精確規劃。撰寫核心主要分為兩個部分:首行的「關鍵影格對齊指令」以及隨後的「三大核心欄位」。針對純文字生影音(T2VA),可以直接由核心欄位開展敘事;但若是涉及首影格(I2VA)、首尾雙影格(FL2VA)或尾影格(L2VA)的任務,則必須在首行明確標註影格在時間軸上的精確秒數(保留至小數點後兩位),使模型能夠在給定圖片的錨定條件下,合理推導視覺演變與動作路徑。

在核心內容的鋪陳上,模型依賴三大區塊來實現聲畫一體的高品質生成。最重要的「多模態整合描述(integrated_multimodal_description)」負責以時間軸為主軸,詳盡定義畫面風格、構圖變化、主體特徵與實體動作。當影片包含鏡頭切換時,第二個鏡頭起必須嚴格依序標註推進時間戳記;同時,鏡頭運動應將運動類型、幅度與速度自然融入敘述中,而非堆疊孤立的標籤詞彙。涉及角色對話或歌唱時,需賦予穩定的角色識別編號(如 S1、S2),並嚴格使用 <d>...</d> 對話標籤進行封裝;標籤內部僅能放置語言標記(如 [English]、[Chinese])與逐字完全一致的台詞原文,發話者的性別、年齡、聲線音色、語速等特徵與肢體反饋則必須留在標籤之外,若遇到台詞跨鏡頭或被片尾截斷時,還可靈活搭配 <scenetrans> 或 <cutoff> 輔助標籤進行銜接。最後輔以「整體環境音場(overall_soundscape)」與「非敘事性背景音樂(non_diegetic_music)」,前者統整環境音、肢體音與動作音效,後者則純粹以樂器配器、節奏與動態變化來描繪觀眾端聽見的配樂,兩者共同建構出層次分明且高擬真度的立體聲體驗。
| 任務模式 | 首行影格對齊指令規範 | 敘事與動作發展核心邏輯 |
|---|---|---|
| T2VA(純文字生影音) | 無需填寫對齊指令,直接進入三大核心欄位 | 完全由文本從零建立視覺與音訊時間線,詳述開場場景、主體動作演變與鏡頭轉移。 |
| I2VA(首影格生影音) | 聲明第一張圖片對齊於 0.00 秒處 | 以首張圖片為起點錨定主體外貌、服裝與構圖,隨後展開動作的發生、連續發展與反應結果。 |
| FL2VA(首尾影格插值) | 聲明第一張圖對齊 0.00 秒,第二張圖對齊結束秒數 | 聚焦於起點至終點之間的動作過渡、姿態演變與場景光影變化,通常以單鏡頭平滑過渡為主。 |
| L2VA(尾影格推導) | 聲明參考圖片對齊於最終結束秒數 | 反向推導前置可能狀態,精確描述主體與鏡頭如何逐步運動並最終收斂著陸至目標畫面。 |
掌握這套語義結構與鏡頭語彙後,創作者能以標準化格式精準引導 MiniMax H3 生成符合預期的影音內容。以下提供四種常見任務模式的官方標準提示詞範例供參考:
1. 純文字生影音(T2VA)範例
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.
overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.
non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.
2. 首影格生影音(I2VA)範例
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.
overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.
non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
3. 首尾影格插值(FL2VA)範例
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.
overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.
non_diegetic_music: N/A
4. 尾影格推導(L2VA)範例
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.
overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.
non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.
《上一篇》ComfyUI x MiniMax H3:生成 Live2D 風格二次元日系動態壁紙 









留言區 / Comments
萌芽論壇