Chuyển động người tốt hơn cho MiniMax H3: chuyển động cơ thể tự nhiên, nhất quán
Chủ đề: Mô hình
Bởi Captain
Đăng ngày 2026-10-01
Ảnh mẫu
Prompt và thiết lập do những người tạo ra các ảnh này trên Civitai chia sẻ. Chọn một ảnh để xem nó được tạo như thế nào.
Thiết lập
Số bước
12
Guidance
1
Kích thước
1440x832
Prompt
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] The scene opens from <Picture 1>, preserving the three young women standing close together outdoors on a city street in winter, with European-style buildings and bare trees in the background under soft daylight. The camera has a natural handheld smartphone selfie shakiness, as if held by one of the girls at arm's length. The girl on the left with long straight black hair, a thick gray scarf, and a dark coat (S1) shivers slightly, pulling her scarf tighter around her neck with one hand. She looks directly into the camera lens with a bright, slightly complaining expression and says: <d>[Chinese] 今天有点冷啊,快看镜头啊!我们三个谁最好看啊?</d> The girl in the middle with chestnut-brown hair, a gray coat with bow buttons, and a white phone with a bear charm (S2) turns her gaze toward the camera, a playful smile forming on her lips. The girl on the right with a single braid, a gray sweater, and a gold necklace (S3) continues looking down at her dark phone, smiling faintly.
[Shot 2] At 00:05.000, the handheld camera wobbles slightly as the girls adjust their positions. The middle girl (S2) raises her hand in a cute peace-sign gesture beside her face, tilts her head playfully, and laughs with a cheeky, confident expression, saying: <d>[Chinese] 当然是我啦!</d> She winks one eye playfully after the line. The left girl (S1) reacts with a mock-offended expression, opening her mouth in playful protest and lightly nudging the middle girl's shoulder with her elbow. The right girl (S3) finally looks up from her phone toward the camera, an amused smile spreading across her face as she shakes her head at the other two.
[Shot 3] At 00:10.000, the right girl (S3) laughs warmly, her eyes crinkling with genuine amusement, and says in a light, scolding-but-affectionate tone: <d>[Chinese] 两个臭屁精!别拍了,赶快吃火锅去啦!</d> She puts her dark phone away into her coat pocket and starts to turn her body toward the street, gesturing with her hand for the others to follow. The middle girl (S2) giggles and lowers her peace-sign hand, while the left girl (S1) laughs and nods in agreement. The handheld camera begins to lower and tilt as if the person filming is putting the phone down, the three girls starting to move together toward the street, their shoulders brushing as they walk off, still laughing. The frame ends with a slight downward tilt as the selfie video concludes.
overall_soundscape:
Soft outdoor city street ambience continues throughout — distant traffic hum, occasional car passing, faint wind rustling through the bare tree branches, and the low murmur of pedestrians in the background. The girls' voices are clear and close, captured by the phone's microphone with a slight natural reverb from the open street. Subtle fabric rustling is audible as they adjust their coats and scarves, and the soft click of a phone lock button is heard near the end as the right girl puts her phone away. Light, cheerful laughter from all three girls overlaps naturally throughout the conversation.
non_diegetic_music:
N/A
Thiết lập
Kích thước
416x736
Prompt
the girl from example.png finally meets ryu from streetfighter in a dark ally. the girl proves far more superior in hand to hand combat since she is fighting in her own comfyui environment using a modified workflow.
Thiết lập
Sampler
Euler
Số bước
8
Seed
1765670026011586
Kích thước
512x800
Prompt
[source image: https://civitai.red/images/141930523]
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the end mark of the target video.
integrated_multimodal_description: [Shot 1] mvmt. Live-action, cinematic, a wide shot captures a bustling European city square under bright daylight, featuring richly textured historic buildings surrounding the open plaza, with the camera remaining static, perfectly framing the architectural grandeur. Overall, the scene maintains a pleasant, cheerful atmosphere underscored by upbeat, lively instrumental music. After a period of stillness, at 00:05.000, a distant, drawn-out female shriek of fright drifts faintly through the music. As the sound of the scream grows noticeably louder, overpowering the music's volume, the broom with a woman flying headfirst downwards streaks at high speed across the frame in half of a second from left to right straightly. As the broom carries her out of the frame to the right, her shriek diminishes in volume, gradually being swallowed again by the music. For several seconds, the frame returns to a static depiction of the city square with no significant action.
[Shot 2] At 00:07.500, Medium low-angle shot of the same woman who is now seated upright the broom, seemingly having regained control, the camera placed just in front of her broom handle. With a happy expression on her face, with her hair and scarf blowing in the air and her eyes and mouth widely opened in cry of mixed joy and fear, mingled with laughter, she rapidly flies over the city plaza while the camera tracking her non-stop flight.
overall_soundscape: The primary ambient sound is the distant murmur of the city crowd. The distinct, loud, and then fading shrieks punctuate the general urban noise. The laugh and cries of joy of the woman when she filmed in [Shot 2].
non_diegetic_music: Upbeat, lively instrumental music plays throughout, maintaining a consistent, cheerful tempo, though it is periodically overwhelmed by the screams.
Thiết lập
Số bước
10
Kích thước
768x1216
Prompt
subject_definitions:
<Subject 1> is the veteran archmage, with his facial identity and clothing perfectly matching <Picture 1>, his voice weathered, gravelly, low, ancient and commanding.
<Picture 1> is the opening-frame anchor for [Shot 1], the archmage atop the cliff's edge analyzing the battlefield below.
<Audio 1> is the original background music track for this video, gradually building in intensity over its full duration.
summary:
[keyframe completion + audio reuse] The target video opens from <Picture 1>, the archmage atop the cliff. Over the full 15 seconds, he delivers a lengthy incantation while a storm gathers behind him and the camera arcs from his front to his back, all set to <Audio 1>, which builds gradually in intensity throughout without needing to align to any specific action.
retention_analysis:
<Picture 1> ([Shot 1] first frame): fully_preserved - the archmage, his staff, his aura, and the cliffside setting anchor the opening composition.
<Subject 1> (appears throughout [Shot 1]): fully_preserved - the archmage's identity, staff, and aura from <Picture 1> are retained throughout.
<Audio 1>: fully_copy - the track is reused in full, unmodified, as the complete non-diegetic score for the entire video, its gradual build not tied to any single visual cue.
detailed_description: Live-action, cinematic. [Shot 1] <Subject 1>, the veteran archmage matching <Picture 1>, remains atop the cliff's edge exactly as established, analyzing the battlefield below, his staff still crackling with stored power visualized as a swirling aura of light, the sky overcast but bright behind him. The camera begins directly in front of the mage, matching the vantage of <Picture 1>, and over the full 15 seconds of this shot, arcs slowly and continuously around him with small amplitude at slow speed, tracing a wide arc until it arrives directly behind him by the shot's end, the mage's position remaining fixed in place throughout while the camera alone circles around him. Below on the battlefield, distant armies clash in a chaotic, evenly matched struggle, neither side gaining ground, soldiers and war machines colliding in a constant, disorganized churn visible far below the cliff. From 00:00.000 to 00:05.000, a strong wind begins gusting from behind the mage, his robes snapping and billowing loudly with each fresh surge, the aura around his staff pulsing brighter in time with the wind's intensity. He (S1) begins his incantation, <d>[Japanese] 天をつらぬく黄金のことわり、我が身に降りよ。</d> A white subtitle reading "Golden law that pierces the heavens, descend upon me." appears at the bottom of the frame, slowly raising his staff overhead with both hands as he speaks, lifting it toward the sky in a deliberate, unhurried motion. From 00:05.000 to 00:09.000, the camera continues its slow arc as the overcast sky behind him churns and thickens, towering cumulonimbus clouds building rapidly upward from the wind's disturbance, darkening progressively as they grow, casting a spreading shadow across the chaotic, struggling battlefield below. He continues, his tone sharpening with authority, <d>[Japanese] 閃光のさばき、時よ止まれ。</d> A white subtitle reading "Judgment of the flash, let time stand still." appears at the bottom of the frame, his staff now held fully raised above his head, steady and unmoving. From 00:09.000 to 00:12.500, still arcing steadily around him, the towering clouds above have grown massive enough to plunge the battlefield into a deep, shifting shadow, static electricity visibly building within the cloud mass, faint flickers of inner light sparking intermittently through the dark clouds. He presses on, his voice rising with intensity, <d>[Japanese] いにしえの盟約、今ここに解き放たん。</d> A white subtitle reading "Ancient covenant, I release you here and now." appears at the bottom of the frame, the staff remaining held high overhead, unmoving, in anticipation of the final strike. From 00:12.500 to 00:15.000, the camera completes its arc, arriving directly behind the mage by the shot's end. As he reaches the incantation's final line, his staff's aura flares to its brightest yet, his robes whipping violently in the now-fierce wind, his expression fierce with concentration as he suddenly swings the staff downward in one sharp, decisive motion, bringing it down in front of himself as he delivers the final words, <d>[Japanese] ——きたれ、らいていの顕現!</d> A white subtitle reading "—Come forth, manifestation of thunder!" appears at the bottom of the frame, the shot ending the instant the incantation concludes and the staff completes its downward swing.
overall_soundscape: A low wind begins as a faint murmur and builds steadily into a powerful, sustained gust, the mage's heavy robes snapping and flapping loudly throughout. The distant, chaotic clatter of clashing armies and war machines rumbles faintly from the battlefield far below throughout the entire shot. Beginning around the midpoint, a deep, rolling rumble of distant thunder builds within the gathering clouds, growing steadily louder and closer, punctuated by faint crackling pops of static discharge as the storm intensifies toward the final line, culminating in a sharp whoosh as the staff swings down.
non_diegetic_music: <Audio 1> plays in full for the entire 15-second duration of the video, exactly as the source file, unmodified, its gradual build in intensity underscoring the mage's incantation and the gathering storm without being tied to any specific moment or action.
Thiết lập
Số bước
30
Kích thước
640x1152
Prompt
Eye-level continuous shot of a contemporary dancer in loose linen clothing executing a fluid, spinning turn on one foot, transitioning smoothly into a low spiral crouch onto the wooden floor, and rising back up into an extended reach. 360-degree torso rotation, continuous arm trailing, clean limb tracking, no morphing or phantom limbs.
Thiết lập
Số bước
8
Kích thước
768x1216
Prompt
subject_definitions:
<Subject 2> is the spiker, the woman in <Picture 1>, with her facial features, clothing, and back number 3 perfectly referenced. She has a quiet, reserved voice.
<Subject 7> is the male coach in <Picture 2>, with his facial features, clothing, and penis proportions perfectly. He has the voice of an overweight man.
<Subject 8> is the Japanese gymnasium <Picture 3>, with its space, equipment, and lighting perfectly referenced.
Summary:
[Reference generation] This scene shows <Subject 2> entering the room where <Subject 8> is located and conversing with <Subject 7>, who is off-screen.
retention_analysis:
<Subject 2> (appears in [Shot 1]): fully_preserved - the woman's facial features, clothing, and back number 3 are retained.
<Subject 7> (appears in [Shot 1]): fully_preserved - the male coach's facial features and clothing and penis size are retained.
<Subject 8> (appears in [Shot 1]): fully_preserved - the gymnasium's space, equipment, and lighting are retained.
detailed_description:
The target video is set in a dimly lit, dusty Japanese gymnasium storage room from <Subject 8>, evoking a sense of unease and isolation.
[Shot 1] a medium wide shot establishes <Subject 8>, the gymnasium storage room, with a dimly lit interior and a dusty atmosphere. Accompanied by a sound like that of a heavy cart being moved, the metal sliding door slowly side opens. The camera begins a slow, ascending low-angle shot of <Subject 2>, the female volleyball team member, in her volleyball uniform. She appears tense and apprehensive. A quiet, unsettling piano melody begins to play softly in the background. Upon entering the room, she closes the sliding door behind her without turning around; the sound produced is like that of moving a heavy cart. <Subject 2> turns toward the camera and <Subject 2> (S1) says in a quiet, reserved voice, <d>[Japanese] コーチ、何かありましたか?</d> A white subtitle reading "Coach, is there something?" appears at the bottom of the frame. <Subject 7> (S2) voice (off-screen, low and calm) says: <d>[Japanese] 3番ちゃん、頑張ってるし、少し体をほぐしてあげようと思ってね</d> A white subtitle reading "3-chan, you’ve been working hard, so I thought I’d loosen up your body a little." appears at the bottom of the frame. <Subject 2> stands frozen, her back rounded with anxiety. Ventilation fan steady, rhythmic hum fills the silence.
overall_soundscape:
Soft indoor gymnasium room tone and a low ventilation hum.
non_diegetic_music:
Ambient noise and low-pitched drone sounds play faintly, accentuating a claustrophobic and oppressive atmosphere. Low-frequency vibrations create an unpleasant auditory sensation.
Tổng quan
Better Human Motion (v1.0-MH3) là một
LoRA
cho các mô hình video
MiniMax H3
và
LTX
giúp cải thiện chuyển động cơ thể: đi bộ mượt mà hơn, cử chỉ và quay người tự nhiên, ít lỗi chi hơn và giải phẫu nhất quán hơn qua các khung hình. Nó được huấn luyện trên text-to-video nhưng cũng dùng được cho image-to-video.
Khuyến nghị của tác giả: 720x1280, 15-30 bước, trọng số 0.4-0.8, và câu ngắn để tránh prompt bị lẫn.
Tham khảo
Tên | Kiểu | Ý nghĩa |
|---|---|---|
LoRA weight | 0.4-0.8 | Mặc định 0.6. |
Resolution | 720x1280 | Video dọc; video ngang cũng được với cùng số pixel. |
Steps | 15-30 | |
Prompt style | short sentences | Mỗi câu một hành động. |
Mode | T2V / I2V | Cả hai. |
AD

Không cần cài đặt – chạy mô hình AI trên đám mây
BitVector Prism là cách dễ nhất để bắt đầu: chọn một mô hình, nhập prompt và tạo ảnh trong một ứng dụng web gọn gàng, không cần cấu hình gì. BitVector cũng có trên Discord và trên web (SpyGlass).
Từng bước
- Viết prompt H3 với một hoặc hai hành động rõ ràng trong câu ngắn.
- 720x1280, 20 bước.
- Giữ âm thanh đơn giản để chuyển động được chú ý.
- Tăng lên 0.8 cho nhảy hoặc thể thao; giảm xuống 0.4 cho cử chỉ nhẹ nhàng.
Ví dụ
Đi bộ
integrated_multimodal_description: [Shot 1] A woman in a trench coat walks along a rainy pier toward the camera. She adjusts her collar. She stops and looks at the sea.
overall_soundscape: Rain, waves, footsteps on wood.
non_diegetic_music: N/A
Better Motion 0.6
Nhảy
integrated_multimodal_description: [Shot 1] A dancer in a studio spins once and lands in a low pose. Mirrors behind. Camera static.
overall_soundscape: Soft shoes on wood floor.
Better Motion 0.8
Ý tưởng sáng tạo: anime gặp phim thực
integrated_multimodal_description: [Shot 1] A young woman in a hoodie faces a martial artist in a dark alley. She steps forward and blocks a slow punch. Camera handheld.
Better Motion 0.7
Mẹo
- Mô tả đường đi chuyển động (đi từ trái sang phải) thay vì dùng tính từ.
- Kết hợp với Combat LoRA ở trọng số thấp hơn cho cảnh hành động.
- Với I2V, đảm bảo khung bắt đầu đã thể hiện tư thế bắt đầu chuyển động.
- Clip ngắn hơn (5 giây) giữ được sự nhất quán hơn.
Khắc phục sự cố
Prompt bị lẫn (hành động trộn lẫn)
Vì sao xảy ra
Câu ghép dài.
Cách sửa
Chia thành câu ngắn, mỗi câu một hành động.
Chuyển động cứng
Vì sao xảy ra
Trọng số quá thấp.
Cách sửa
0.7-0.8.
AD

Không cần cài đặt – chạy mô hình AI trên đám mây
PirateDiffusion chỉ có trên Telegram và dành cho người dùng chuyên nghiệp: hàng nghìn mô hình, LoRA và workflow điều khiển bằng lệnh chat, tạo không giới hạn với gói giá cố định.
Câu hỏi
LTX cũng được không?
Có, có phiên bản cho LTX.
Có trigger không?
Không có.
Liên kết và nguồn
Mô hình trong hướng dẫn này
Người viết
Captain
Hướng dẫn liên quan
Mô hình
MiniMax H3 Combat BASE V2: chuyển động chiến đấu, va chạm và kịch tính
LoRA Combat BASE của FourBunny cho MiniMax H3 cải thiện biên đạo chiến đấu, phản ứng khi bị đánh, tính liên tục và nguyên nhân kết quả vật lý, đồng thời cũng giúp các cảnh đối thoại. Tìm hiểu các trigger prfight1, prfin1 và prslow1 và cách dựng cảnh.
Bởi Captain
2026-10-01
Mô hình
GalaxyAce LoRA cho MiniMax H3: video điện thoại tầm thấp đầu những năm 2010
GalaxyAce LoRA của aiguild tái tạo lại kiểu ảnh của camera điện thoại Samsung Galaxy Ace: màu sắc nhạt, nhiễu cảm biến, hiệu ứng cuộn màn trập và cách khung hình đời thường. Tìm hiểu điểm mạnh theo từng model, prompt và ý tưởng video tìm thấy.
Bởi Captain
2026-10-02
Mô hình
MiniMax H3 Turbo LoRAs: Video 4-8 bước với gói larryvrh
Gói MiniMax H3 Turbo LoRA (larryvrh 4 bước và bạn bè) giảm số bước tạo video từ hàng chục xuống chỉ còn vài bước. Tìm hiểu số bước, bộ lấy mẫu và độ mạnh từng biến thể, cùng cách viết prompt cho H3 để có clip nhanh và sạch.
Bởi Captain
2026-10-01
Mô hình
LoRA Hôn Nồng Nàn cho MiniMax H3: cảnh lãng mạn thuyết phục giữa người lớn
LoRA của kermitfrog1202 cải thiện các cảnh hôn trong video MiniMax H3. Tìm hiểu câu kích hoạt, các giá trị cường độ tác giả tìm ra cho các cặp đôi khác nhau và cách dựng cảnh lãng mạn tinh tế.
Bởi Captain
2026-10-01
Viết prompt
Hướng dẫn MiniMax H3: từ văn bản thành video, từ hình ảnh thành video, video tham khảo và âm thanh
Cách điều khiển MiniMax H3 qua lệnh chat: công thức WHO + WHERE + ACTION + CAMERA + SOUND cho việc tạo video từ văn bản, làm chuyển động cho ảnh tĩnh với hình ảnh thành video, giao việc cho từng ảnh tham khảo trong chế độ tham khảo, và tạo thoại rõ ràng thay vì lí nhí. Có kèm các prompt để sao chép và thử, cùng PDF của tác giả.
Bởi Mark White
2026-09-20
Hướng dẫn khác
Dịch vụ đám mây
Phí cố định so với token: tại sao các gói không giới hạn như Graydient.ai là lựa chọn tốt nhất cho nhà sáng tạo AI năm 2026
So sánh chi phí có xếp hạng, với giá được kiểm tra vào ngày 8 tháng 10 năm 2026, giữa các gói không giới hạn phí cố định như Graydient.ai với giá token và tín dụng từ Midjourney, Leonardo, fal.ai, Replicate, Runway và Kling: chi phí thực sự cho 1.000 hình ảnh và 1.000 video mỗi tháng trên mỗi dịch vụ, và tại sao một khoản phí cố định cho hình ảnh, video, âm thanh, trò chuyện Grok LLM và ứng dụng web không giới hạn là giá trị tốt nhất cho những người sáng tạo hay thử nghiệm.
Bởi Captain
2026-10-08
Mô hình
FLUX.1 Dev: cách tạo prompt, cài đặt tốt nhất và ý tưởng sáng tạo
Hướng dẫn thực tế về FLUX.1 Dev của Black Forest Labs: cách tạo prompt bằng ngôn ngữ tự nhiên, cài đặt hướng dẫn và bước, xếp chồng LoRA, hiển thị chữ và ý tưởng prompt tận dụng điểm mạnh của nó.
Bởi Captain
2026-09-30
Mô hình
Stable Diffusion XL 1.0: lời nhắc, cài đặt và những gì nó vẫn làm tốt nhất
Cách tận dụng tối đa mô hình cơ sở chính thức SDXL 1.0: độ phân giải gốc, cài đặt CFG và sampler, bộ tinh chỉnh, lời nhắc tiêu cực và ý tưởng phong cách mà SDXL vẫn nổi bật.
Bởi Captain
2026-10-01
Mô hình
Stable Diffusion 1.5: mô hình cổ điển, dùng đúng cách
Stable Diffusion 1.5 vẫn đáng để biết: độ phân giải phù hợp, CFG và sampler, cách dùng thư viện lớn LoRA và embedding của nó, cùng ý tưởng prompt sáng tạo phù hợp với mô hình 512 pixel.
Bởi Captain
2026-10-02
Mô hình
FLUX.2 Dev: hướng dẫn sử dụng model mới nhất của Black Forest Labs
FLUX.2 Dev hỗ trợ prompt dài hơn, văn bản tốt hơn, thực tế hơn và chỉnh sửa đa tham chiếu. Hướng dẫn này bao gồm quy trình turbo, cài đặt hướng dẫn và độ phân giải, cấu trúc prompt và ý tưởng tận dụng điểm mạnh mới của nó.
Bởi Captain
2026-10-03
Mô hình
Qwen-Image: prompt dài, chữ chuẩn và poster song ngữ
Qwen-Image là model bạn nên dùng khi chữ trong ảnh quan trọng. Học cách viết prompt mô tả dài, tạo chữ tiếng Anh và Trung chính xác, chỉnh CFG và bước, và khám phá ý tưởng sáng tạo với bố cục phức tạp.
Bởi Captain
2026-09-29

Model
Trends
.ai
© ModelTrends.ai
|
© 2026
|
Mọi quyền được bảo lưu





