Model
Trends
.ai
Model
Trends
.ai
the best open source ai models
Autoplay
Run Locally
ltx2-eat-v2.safetensors
1.03 GB · SafeTensor · bf16
Downloads are served by Civitai. Some creators ask for a Civitai login first.
About This Model
LTX2 · LoRAs
Created by
kabachuha
on Civitai
by kabachuha - Cartoonishly Eaten - LTX-2 / Wan2.1 14b T2V
Creator notes
LTX-2 Eat LoRA that gobbles everyone! (Now with sound (really, turn it on on examples)) Basically, you start the video with a subject. Then, suddenly, the camera zooms out revealing that the subject is now miniaturized and then another character steps up and eats them cartoonishly and non-graphically. This is my fifth LTX-2 LoRA (published globally). Now, this is the start of porting my legacy Civitai loras from Wan to LTX-2. This LoRA is best working with first-last frame, however start frame may be sufficient if you describe the other subject well. Beware, FLF inherits all LTX-2's flaws and it can do slideshow-like things from time to time (Idk why it spawns the first and the end frame at the end, best way is to simply cut it). Easter egg: characters can devour themselves in a loop if you set the first and the last frames the same pictures. In contrast to all my previous LTX-2 LoRAs, this one was superhard to train. With CREPA, TREAD, FFN unfreeze, higher rank, Prodigy, the loss didn't lower much and even showed signs of divergence (initially stable loss curve eventually progressing to insanely frequent oscillations without decline). Needless to say, all I could see was pure body horror. With the tongues, the hands themselves being eaten, distorted limbs, etc. Then I remembered that for sharper results in the case of high oscillations not MSE, but Huber loss is needed. I used scheduled Huber loss (exponential) and it much stabilized the loss curve, producing the much needed downturn at last. Interestingly, this loss choice caused the CREPA regularization loss curve's shape not be just a monotonous sigmoid and even have smooth hills. Warning: because deep features CREPA or TREAD was used, some of the videos might have slightly washed out feel. If you experience it, try adding vivid colors to the positive prompt, and things like washed out,gray to the negative prompt, and also if the start images are themselves vivid, it will go much better. The runtime for this experiment totaled 5 hours (and five failed attempts, ranging up to 8 hours). The hardware used for training was 1x5090, with zero blocks swapped, ~4 s/it. The dataset consists of 6 organic; video fragments (repeated 2 times), which the original LoRA was trained on, plus 47 picked Wan2.2 generations made with that LoRA applied. Overall, the final checkpoint was picked at 4000 optimization steps. The SimpleTuner training and dataset configs are under config.json and ltx2-multiresolution-eat-t2v-v2.json respectively. The ComfyUI workflows are inside the .mp4 video files or on the Huggingface repo. The Huggingface host for the LTX-2 LoRA is at . Trigger words You should use eat style to trigger the image generation. Actually, you shouldn't now, because LTX-2 will add an utterance "eat style" at the beginning of the video. Just describe the action similar to the prompts from the examples and it will do the job! For Wan2.1 (legacy): This LoRA introduces the concept of subjects and things eaten in cartoon-like way by being suddenly tossed into a giant mouth and then chewed and consumed non-graphically. To choose the eater and/or the thing being eaten, use VACE. The examples illustrate each mode, in order: 2 x first-frame2video, first-last2video, last-frame2video, pure text2video (not recommended, as it's slightly retrained in this mode). The generation resolution is advised to be 512x512 or close in the spatial dimensions, with recommended duration 45 base frames (49 total, 3 seconds). The t2v training was made using diffusion-pipe for 100 epochs and flow shift of 4.5. For image2video/flf2video recommended to use with the kijai VACE workflows, standard 1.0 lora weight, 4.0-6.0 cfg, 8.0-16.0 shift + cfg_zero_star. (see videos meta in comfy) Best works on cartoon-stylized and anime characters, can be weird on realistic. For realistic, supplying an additional existing cartoon-style reference is advised. Known issues: the object sometimes is bit/chewed like a gum, but not swallowed. (can be slightly countered with adding pushing it inside with the hand.) The trigger word is 'eat style'. The best prompts are: """ eat style. The video begins with [object]. Then a gigantic cartoon hand seizes the [object] from below and tosses it into a gigantic mouth, which appeared on the right side. The camera zooms out, showing the new [eater] chewing and fully swallowing the old tiny [object]. """ P.S. In case you can't find the metadata, the example workflow for the wolf is here
Trigger words
eat style
237 downloads · 15 likes on Civitai
SHA256
0B1D8E28A03492C2986331C0BEE94EBC2A54B41891C85CA772600692EE973185
ModelTrends.ai Model ID
#21105
Flag content
Example renders
Prompts and settings shared by the people who made these renders on Civitai. Pick a render to see how it was made.
Settings
Sampler
Euler
Steps
20
Guidance
4
Seed
32
Size
960x1024
Prompt
Vivid colors. The video begins with a cartoon rabbit. Suddenly, a comically long prehensile sticky anime girl's tongue appears from out of the frame and wraps around the rabbit from all sides, scooping it up. The tongue forms the perfect circle. Then the anime girl's tongue seizes the cartoon rabbit from below and tosses it into a gigantic mouth, which appeared to the right. The anime girl's two arms rest on the ground statically, she only shoots her tongue like a frog. The camera zooms out, showing the new anime girl chewing and fully swallowing the cartoon rabbit. Wet slurping sounds.
Negative prompt
Three arms. She reaches for the rabbit with her hands, the hand touches the rabbit, hand in the frame, washed out, gray, blurry, out of focus, overexposed, underexposed, low contrast, washed out colors, excessive noise, grainy texture, poor lighting, flickering, motion blur, distorted proportions, unnatural skin tones, deformed facial features, asymmetrical face, missing facial features, extra limbs, disfigured hands, wrong hand count, artifacts around text, unreadable text on shirt or hat, incorrect lettering on cap (“PNTR”), incorrect t-shirt slogan (“JUST DO IT”), missing microphone, misplaced microphone, inconsistent perspective, camera shake, incorrect depth of field, background too sharp, background clutter, distracting reflections, harsh shadows, inconsistent lighting direction, color banding, cartoonish rendering, 3D CGI look, unrealistic materials, uncanny valley effect, incorrect ethnicity, wrong gender, exaggerated expressions, smiling, laughing, exaggerated sadness, wrong gaze direction, eyes looking at camera, mismatched lip sync, silent or muted audio, distorted voice, robotic voice, echo, background noise, off-sync audio, missing sniff sounds, incorrect dialogue, added dialogue, repetitive speech, jittery movement, awkward pauses, incorrect timing, unnatural transitions, inconsistent framing, tilted camera, missing door or shelves, missing shallow depth of field, flat lighting, inconsistent tone, cinematic oversaturation, stylized filters, or AI artifacts.
Settings
Sampler
Euler
Steps
20
Guidance
4
Seed
29
Size
1024x1024
Prompt
Vivid colors. The video begins with a cartoon girl blowing a chewing gum balloon. Then a gigantic anime girl hand seizes the cartoon girl from below and tosses her into a gigantic mouth, which appeared to the right. The camera zooms out, showing the new anime girl chewing and fully swallowing the cartoon girl. Wet slurping sounds.
Negative prompt
washed out, gray, blurry, out of focus, overexposed, underexposed, low contrast, washed out colors, excessive noise, grainy texture, poor lighting, flickering, motion blur, distorted proportions, unnatural skin tones, deformed facial features, asymmetrical face, missing facial features, extra limbs, disfigured hands, wrong hand count, artifacts around text, unreadable text on shirt or hat, incorrect lettering on cap (“PNTR”), incorrect t-shirt slogan (“JUST DO IT”), missing microphone, misplaced microphone, inconsistent perspective, camera shake, incorrect depth of field, background too sharp, background clutter, distracting reflections, harsh shadows, inconsistent lighting direction, color banding, cartoonish rendering, 3D CGI look, unrealistic materials, uncanny valley effect, incorrect ethnicity, wrong gender, exaggerated expressions, smiling, laughing, exaggerated sadness, wrong gaze direction, eyes looking at camera, mismatched lip sync, silent or muted audio, distorted voice, robotic voice, echo, background noise, off-sync audio, missing sniff sounds, incorrect dialogue, added dialogue, repetitive speech, jittery movement, awkward pauses, incorrect timing, unnatural transitions, inconsistent framing, tilted camera, missing door or shelves, missing shallow depth of field, flat lighting, inconsistent tone, cinematic oversaturation, stylized filters, or AI artifacts.
Settings
Sampler
Euler
Steps
20
Guidance
4
Seed
26
Size
1024x1024
Prompt
The video begins with a realistic tank. Then a gigantic cartoon cat maw hand seizes the tank from below and tosses it into a gigantic mouth, which appeared to the right. The camera zooms out, showing the new cartoon cat chewing, tearing apart and fully swallowing the realistic tank.
Negative prompt
blurry, out of focus, overexposed, underexposed, low contrast, washed out colors, excessive noise, grainy texture, poor lighting, flickering, motion blur, distorted proportions, unnatural skin tones, deformed facial features, asymmetrical face, missing facial features, extra limbs, disfigured hands, wrong hand count, artifacts around text, unreadable text on shirt or hat, incorrect lettering on cap (“PNTR”), incorrect t-shirt slogan (“JUST DO IT”), missing microphone, misplaced microphone, inconsistent perspective, camera shake, incorrect depth of field, background too sharp, background clutter, distracting reflections, harsh shadows, inconsistent lighting direction, color banding, cartoonish rendering, 3D CGI look, unrealistic materials, uncanny valley effect, incorrect ethnicity, wrong gender, exaggerated expressions, smiling, laughing, exaggerated sadness, wrong gaze direction, eyes looking at camera, mismatched lip sync, silent or muted audio, distorted voice, robotic voice, echo, background noise, off-sync audio, missing sniff sounds, incorrect dialogue, added dialogue, repetitive speech, jittery movement, awkward pauses, incorrect timing, unnatural transitions, inconsistent framing, tilted camera, missing door or shelves, missing shallow depth of field, flat lighting, inconsistent tone, cinematic oversaturation, stylized filters, or AI artifacts.
Settings
Sampler
Euler
Steps
20
Guidance
4
Seed
31
Size
1024x1024
Prompt
Vivid colors. The video starts with a close-up of a cute anime girl. Suddenly, a colossal same-looking anime girl's hand appears on the right, swoops down and scoops her up. The colossal anime girl looks exactly like the first girl. The camera pans up to reveal the same-looking colossal girl's gaping maw, filled with rows of many sharp teeth. The tiny anime girl is tossed directly into the mouth and disappears instantly, swallowed whole. The camera zooms out. The same-looking colossal girl lets out a satisfied moan. The camera zooms out even more with now the new same looking girl standing in place of the original girl, with the same pose and the animation forms a perfect loop.
Negative prompt
washed out, gray, blurry, out of focus, overexposed, underexposed, low contrast, washed out colors, excessive noise, grainy texture, poor lighting, flickering, motion blur, distorted proportions, unnatural skin tones, deformed facial features, asymmetrical face, missing facial features, extra limbs, disfigured hands, wrong hand count, artifacts around text, unreadable text on shirt or hat, incorrect lettering on cap (“PNTR”), incorrect t-shirt slogan (“JUST DO IT”), missing microphone, misplaced microphone, inconsistent perspective, camera shake, incorrect depth of field, background too sharp, background clutter, distracting reflections, harsh shadows, inconsistent lighting direction, color banding, cartoonish rendering, 3D CGI look, unrealistic materials, uncanny valley effect, incorrect ethnicity, wrong gender, exaggerated expressions, smiling, laughing, exaggerated sadness, wrong gaze direction, eyes looking at camera, mismatched lip sync, silent or muted audio, distorted voice, robotic voice, echo, background noise, off-sync audio, missing sniff sounds, incorrect dialogue, added dialogue, repetitive speech, jittery movement, awkward pauses, incorrect timing, unnatural transitions, inconsistent framing, tilted camera, missing door or shelves, missing shallow depth of field, flat lighting, inconsistent tone, cinematic oversaturation, stylized filters, or AI artifacts.
Settings
Sampler
Euler
Steps
20
Guidance
4
Seed
27
Size
1024x1024
Prompt
Vivid colors. The video starts with a close-up of two cute anime girls. Suddenly, the small anime girl in the right, stretches hands and scoops the left girl up. The camera pans up to reveal the small girl's gaping maw, filled with rows of many sharp teeth. The anime girl on the left is tossed directly into the right girl's mouth and disappears instantly, swallowed whole. The camera zooms out. The smaller girl on the right lets out a satisfied moan, her breasts jiggling as she chews. In the end only the girl on the right is left in the frame.
Negative prompt
blurry, out of focus, overexposed, underexposed, low contrast, washed out colors, excessive noise, grainy texture, poor lighting, flickering, motion blur, distorted proportions, unnatural skin tones, deformed facial features, asymmetrical face, missing facial features, extra limbs, disfigured hands, wrong hand count, artifacts around text, unreadable text on shirt or hat, incorrect lettering on cap (“PNTR”), incorrect t-shirt slogan (“JUST DO IT”), missing microphone, misplaced microphone, inconsistent perspective, camera shake, incorrect depth of field, background too sharp, background clutter, distracting reflections, harsh shadows, inconsistent lighting direction, color banding, cartoonish rendering, 3D CGI look, unrealistic materials, uncanny valley effect, incorrect ethnicity, wrong gender, exaggerated expressions, smiling, laughing, exaggerated sadness, wrong gaze direction, eyes looking at camera, mismatched lip sync, silent or muted audio, distorted voice, robotic voice, echo, background noise, off-sync audio, missing sniff sounds, incorrect dialogue, added dialogue, repetitive speech, jittery movement, awkward pauses, incorrect timing, unnatural transitions, inconsistent framing, tilted camera, missing door or shelves, missing shallow depth of field, flat lighting, inconsistent tone, cinematic oversaturation, stylized filters, or AI artifacts.
Settings
Sampler
Euler
Steps
20
Guidance
4
Seed
31
Size
1024x1024
Prompt
Vivid colors. The video begins with a playful anime girl who is waving her hand. Then a gigantic anime girl to the right behind sucks the air like a vacuum cleaner. The bigger anime girl's hands stand still on her waist and she uses her mouth alone. The air pulls the smaller anime girl into a gigantic mouth, which appeared to the right. The camera zooms out, showing the new anime girl chewing and fully swallowing the cartoon girl. Wet slurping sounds, piano music. Then the camera zooms out.
Negative prompt
washed out, gray, blurry, out of focus, overexposed, underexposed, low contrast, washed out colors, excessive noise, grainy texture, poor lighting, flickering, motion blur, distorted proportions, unnatural skin tones, deformed facial features, asymmetrical face, missing facial features, extra limbs, disfigured hands, wrong hand count, artifacts around text, unreadable text on shirt or hat, incorrect lettering on cap (“PNTR”), incorrect t-shirt slogan (“JUST DO IT”), missing microphone, misplaced microphone, inconsistent perspective, camera shake, incorrect depth of field, background too sharp, background clutter, distracting reflections, harsh shadows, inconsistent lighting direction, color banding, cartoonish rendering, 3D CGI look, unrealistic materials, uncanny valley effect, incorrect ethnicity, wrong gender, exaggerated expressions, smiling, laughing, exaggerated sadness, wrong gaze direction, eyes looking at camera, mismatched lip sync, silent or muted audio, distorted voice, robotic voice, echo, background noise, off-sync audio, missing sniff sounds, incorrect dialogue, added dialogue, repetitive speech, jittery movement, awkward pauses, incorrect timing, unnatural transitions, inconsistent framing, tilted camera, missing door or shelves, missing shallow depth of field, flat lighting, inconsistent tone, cinematic oversaturation, stylized filters, or AI artifacts.
Badges for creators
Made this model? Put its standing on your Civitai page, your Hugging Face card, your README or your site. Each badge is a small picture that shows the standing on the day you copy it and never changes afterwards, however the rankings move later. The date is part of the badge.
cartoonishly eaten on ModelTrends.ai: #52 of 80 Ltx2 models
cartoonishly eaten on ModelTrends.ai: #52 of 67 Ltx2 LoRAs
AI Models for Image, Video & Text Generation
cartoonishly eaten ltx2 is a LoRA for the Ltx2 family listed on ModelTrends.ai, a read-only catalog of open source AI image, video and text models. Compare it with other Ltx2 models, check its heat score to see how it is trending, and open it on PirateDiffusion or BitVector to try it.
Browse by Model Family
Browse all models