选一个模型,输入你的提示词,直接查看生成结果。
HiAPI Blog
HiAPI
现在就用 HiAPI 生成
A hand lifts a plain grey sun-protection jacket off a tote bag on a subway-station bench — nothing unusual there. What makes the 10-second clip work is the flat 2D cartoon sun sitting on top of the footage like a sticker: it starts the clip bouncing and overheated, panics when the jacket's UPF fabric gets shown off, and by the final beat it's cooled all the way down to pale blue, hugging a 2D ice cube against the jacket's hood with its eyes closed. The whole gag is built on a dual-reference-image prompt structure for Seedance 2.0 Ext, and it's worth breaking down because the technique generalizes to almost any "real product + cartoon mascot reaction" ad.
This is the prompt we used to generate the clip, translated in full (the original run was in Chinese; nothing below is simplified or summarized):
Scene: Real-life, handheld-style footage. A hand holds up a light-grey UPF
sun-protection jacket above a black tote bag resting on a bench, in front of
a city subway station entrance. Natural daylight, authentic street
atmosphere, no stylization on the live-action plate itself.
Composite in a single 2D flat-cartoon sun character (see @image1 for the
exact character design — do not alter its proportions, palette, or line
weight). The character is pasted onto the live-action plate like a sticker:
it must stay perfectly flat 2D at all times, ignore the real-world lighting
and shadows around it, and never pick up any 3D shading, reflections, or
perspective distortion from the live footage.
@image1 — locks ONLY the character's product design (the sun icon's shape,
face, color stages, proportions). Do not pull scene, background, hand, or
jacket details from this reference.
@image2 — locks ONLY the live-action scene and the way the jacket is
worn/held (hand position, bench, subway entrance, camera framing). Do not
pull character design from this reference.
Timing (4 comedic beats, synced to the real hand movements in the plate):
1. 0-2s: the sun character appears bright yellow-orange, bouncing with
exaggerated "hot" energy, sitting on top of the folded jacket.
2. 3-5s: the hand lifts the jacket to show the fabric; the character reacts
with mock alarm, sprouts comic sweat drips, color shifting toward
orange-red.
3. 6-8s: the hand drapes the jacket fully open toward camera; the character
visibly cools, color fading from red to pale blue, movement slowing down.
4. 8-10s: climax/reveal - the character, now fully pale blue, hugs a 2D flat
ice cube (same flat-sticker rules apply to the ice cube) and nestles
contentedly against the jacket's hood, eyes closed, totally chilled out.
Sound design: a light comedic "boing" on the character's first bounce, a
cartoon sizzle/steam hiss during the sweating beat, a soft "ahh" exhale plus
a gentle chime on the final cooldown hug. Keep real-world ambient
subway/street sound underneath throughout - don't mute the live-action
audio.
Avoid: no readable text, price tags, UI elements, captions, subtitles, or
watermarks anywhere in the frame. No Chinese text or other text from
@image1/@image2 should leak into the final composite. No 3D shading or
re-lighting on the sticker character at any point.
Two reference images, two completely different jobs. @image1 is a clean turnaround of the sun character art — nothing else. @image2 is a plate (or plate-like reference) of the real scene: the hand, the jacket, the bench, the subway entrance. Splitting them this way stops the model from "contaminating" the cartoon with scene lighting, or dragging scene objects into the character reference. On the API side this maps directly onto two entries in Seedance 2.0 Ext's reference_image_urls array — the prompt's @image1/@image2 placeholders are just 1-indexed positions into that array (reference_image_urls[0], reference_image_urls[1]).
The sticker rule is the whole joke. Telling the model to keep the character "perfectly flat 2D" and to never let it pick up real-world shading is what sells the gag — the character reads as a sticker someone slapped onto real footage, not a CG object inserted into the scene. Drop that constraint and the character starts picking up ambient light and shadow from the subway entrance, and the bit stops being funny.
Comedic timing is tied to real, physical beats in the footage, not to arbitrary timestamps: bounce → show-the-fabric → cool-down → hug-the-ice-cube. Because each beat is anchored to something the hand is actually doing in the plate, the cartoon reaction reads as a direct response to the product rather than a decoration running in parallel.
The negative list is doing real work. "No readable text/price/UI/captions/watermarks" and "no text from the reference images should leak into the final composite" are both there because reference images this specific (studio product shots, street photography) often carry stray signage, price stickers, or packaging text that the model will otherwise happily composite into the final frame.
Seedance 2.0 Ext's schema is strict (unknown fields are rejected outright), so here's exactly what a request needs — prompt, aspect_ratio, and duration are required; resolution, reference_image_urls, and generate_audio are optional:
curl -X POST https://api.hiapi.ai/v1/tasks \
-H "Authorization: Bearer $HIAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2.0@ext",
"input": {
"prompt": "<the full prompt above>",
"aspect_ratio": "9:16",
"duration": 10,
"resolution": "1080p",
"reference_image_urls": [
"https://your-cdn.example.com/sun-character-turnaround.jpg",
"https://your-cdn.example.com/jacket-scene-reference.jpg"
],
"generate_audio": true
}
}'
Poll the returned taskId until it reaches a terminal state:
curl -s https://api.hiapi.ai/v1/tasks/{taskId} \
-H "Authorization: Bearer $HIAPI_API_KEY"
reference_image_urls tops out at 9 entries — this clip only needed 2. At the 1080p tier, Seedance 2.0 Ext bills $0.373 per second, so a 10-second clip like this one comes out to $3.73 (current rates always on the pricing page).
What's the difference between Seedance 2.0's default route and the @ext route?
@ext is a separate route on the same model (seedance-2.0@ext vs. plain seedance-2.0), with its own schema and its own pricing tiers. The ext route is the one that accepts reference_image_urls, which is what makes the dual-reference trick in this prompt possible. See the Seedance 2.0 docs for the full field-by-field breakdown.
Do I need exactly two reference images?
No — the field accepts up to 9. Two was the right number for this prompt because there were exactly two things to lock independently (character design, scene/wearing style). If your mascot needs a third reference (say, a specific pose sheet), add a third entry and reference it as @image3 in the prompt.
Why doesn't the sticker character ever look "lit" by the scene? Because the prompt explicitly forbids it. Without that instruction, the model will often try to be "helpful" and blend the character into the scene's lighting, which kills the sticker effect. State the constraint directly rather than assuming the model will preserve a flat style on its own.
Can I use this with product photos I already have?
Yes — that's the point of splitting the references. Use a clean shot or turnaround of your mascot/character as @image1, and a real photo or video frame of your product in its natural context as @image2. The model composites the former onto the latter.
Does generate_audio cost extra?
No separate line item for it in the current pricing — the per-second rate above already reflects the 1080p tier whether or not audio is generated. Always re-check live pricing before you budget a batch run, since tiers do change.
If you want to try this style of composite on your own product shots, the Seedance 2.0 Ext model page has the current schema and pricing, and the quickstart docs will get you from API key to first task in a few minutes.