在 AI 影像與動畫生成技術持續演進的 2026 年,創作者對於動畫角色動態控制與素材提取的要求日益提升。特別是在製作二次元角色動畫、提取去背遮罩或建構 LoRA 訓練數據集時,「純白背景」配合「連續多樣化動作」是極具實用價值的生成模式。然而,傳統影片生成模型在處理大幅度肢體切換時,容易出現肢體融化、結構崩壞或背景雜訊滋生的問題。MiniMax H3 作為當前開源陣營中頂級的全模態影音生成模型,其強大的提示詞解析與時間軸控制力在此類嚴苛測試中展現了極高的研究價值。
本文將使用我們先前建立的 ComfyUI 雙重加速工作流(整合 Spectrum 頻域快取、SageAttention 注意力機制與 Preview Override 即時預覽),實測 MiniMax H3 在「純白背景下生成角色連續動作動畫」的極限表現。本次測試以本站看板娘「萌芽娘」作為主角,採用 1024 x 1024 px(1:1 正方形、1.0 MP、Multiple = 32)的規格,挑戰在單一鏡頭 10 秒內,以每秒為單位連續切換 10 個完全不同的肢體動作,並同步搭配 120 BPM 的電子節拍音樂與細微動作音效!

▲ 此圖展示了在 ComfyUI 中執行白底角色連續動作動畫生成的完整工作流截圖。左側設定 1024 x 1024 px 正方形解析度並載入萌芽娘立繪,透過雙加速節點驅動 MiniMax H3 進行 10 秒 24 fps 的長序列採樣,右側則是完成影音自動封裝輸出。若需要取得本篇所使用的加速工作流架構,歡迎參考本站先前的工作流教學文章。
📝 Prompt 架構與設計依據解析
為了讓 MiniMax H3 能夠在 10 秒內精確依序執行 10 個不同的動作指令,同時維持純淨無雜質的白底環境,提示詞嚴格依照官方規範進行了深度結構化設計:
- 任務類型與頭部宣告(I2VA):首行標準宣告
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.,確保模型嚴格鎖定傳入的萌芽娘立繪作為 0.00 秒起始幀基準。 - 畫風與純白背景鎖定:指定
2D-animated, a full-body shot並明確要求solid, completely blank white void backdrop,杜絕背景生成多餘的光影雜物,非常適合後續提取乾淨人物遮罩或製作訓練素材。 - 運鏡與鏡頭穩定性:採用標準運鏡語法
The camera holds a static shot throughout the entire duration.,避免鏡頭縮放或晃動造成截圖時人物比例失真。 - 時間戳切分與動作豐富度:在單一鏡頭
[Shot 1]內,以每秒為單位(00:01.000至00:10.000)安排揮手、指向鏡頭、叉腰、抱胸、鞠躬、戰鬥姿態、合十、比 V、側身跑步與單手展示共 10 組姿勢轉換,提供充足的多樣性動態。 - 環境音效與節奏配樂:
overall_soundscape聚焦於肢體動作伴隨的布料摩擦與腳步微動;non_diegetic_music則配置 120 BPM 的電子節拍音樂,遵循官方規範描述樂器、節奏與動態。
📄 完整提示詞內容
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] 2D-animated, a full-body shot against a solid, completely blank white void backdrop, fully referencing the anime character from <Picture 1>. The character begins in the initial standing posture established by <Picture 1>. The camera holds a static shot throughout the entire duration. At 00:01.000, the character smoothly shifts into a friendly waving pose with the right hand raised. At 00:02.000, the character points an index finger directly forward toward the camera. At 00:03.000, the character places both hands firmly on hips in an assertive posture. At 00:04.000, the character crosses both arms across the chest with a slight head tilt. At 00:05.000, the character performs a polite forward bow with arms resting neatly at the sides. At 00:06.000, the character transitions into a dynamic combat-ready stance with one raised fist and one open defensive palm. At 00:07.000, the character brings both hands together in front of the chest in a thoughtful clasping gesture. At 00:08.000, the character flashes a cheerful peace sign next to the cheek. At 00:09.000, the character pivots sideways into a profile running silhouette pose. By 00:10.000, the character settles into a confident standing presentation pose with one open hand extending outward, maintaining a completely isolated figure against the pristine blank background.
overall_soundscape: Light fabric rustling, subtle clothing swishes, and gentle foot shifts accompany the rapid sequence of pose transitions against the otherwise silent background.
non_diegetic_music: A steady rhythmic electronic drum beat at a moderate 120 BPM tempo, layered with crisp synthetic percussion and subtle bass pulses throughout.
🎬 實測連續動作動畫成果展示
▲ 白底角色連續動作動畫生成成果。萌芽娘在純白背景中,精準地按照時間戳記連續流暢變換 10 種不同動作,背景全程維持純淨空白無噪點,同時動作音效與 120 BPM 電子節拍完美對齊!
💡 實測結論
實測結果顯示,MiniMax H3 擁有相當驚人的時間軸語意遵循度(Prompt Following)。只要在提示詞中精確運用時間戳(如 00:01.000 等),模型便能將多個連續複雜的肢體變換平滑串接,且在純白無背景特徵輔助的情況下,人物服裝結構與角色特徵依然維持高度穩定。這套方法非常適合用來批量生成角色動態展示、動畫分鏡參考或是提取高品質乾淨訓練集,展現了強悍的開源生產力價值!😁
《上一篇》【2026 年小房間整修工程】第十八集:廁所復原、房間牆面恢復 









留言區 / Comments
萌芽論壇