Most people prompting a continuous narrative video clip run into the same failure mode: the opening frame matches what they want, but by second eight the model has drifted — different robot proportions, a different mist color, a camera move that doesn't land where the story needs it to. Seedance 2.0 Ext solves this with two distinct anchoring strategies, and picking the wrong one for your use case is the difference between a clean 12-second shot and three wasted re-runs.
Both techniques below come from the same narrative world — a rain-soaked mountain city, an orange-and-white monorail, and a compact maintenance robot repairing a damaged roof connector — run on the exact same model, seedance-2.0@ext. The only thing that changes is which frames you anchor.
The same story, two different prompting strategies
First-frame anchoring gives the model one fixed starting point (a reference image) and lets it improvise everything after that — camera movement, pacing, and where the shot ends up. First-last-frame interpolation gives it two fixed points, start and end, and the model fills in a continuous path between them. The prompts, sound design, and subject continuity are nearly identical in both clips below; the control mechanism is what's different.
Technique 1: First-frame hard anchoring
This clip anchors only the opening frame, then lets the camera move freely through a single uninterrupted tracking shot — rail-height, to train-roof level, to a pulled-back reveal of the city.
Use the supplied image as the exact first frame and hard visual anchor: preserve the same original orange-and-white maintenance robot, the same rainy mountain city, the same monorail geometry, wet rail material, amber train lights, and blue mist. 12-second vertical cinematic action film, one uninterrupted continuous tracking shot with no cuts. 0-3s: start at the exact supplied frame; camera is inches above rain-soaked rail tracks, then accelerates forward and slightly upward toward the train roof. Puddles ripple from rail vibration and rain strikes the lens naturally. 3-7s: track beside the moving monorail roof; the compact maintenance robot braces into the wind and pulls the loose heavy orange high-voltage cable toward the damaged roof connector. Show realistic cable weight, tension, metal friction, wheel vibration and water spray. 7-10s: a brief blue electrical arc snaps across the gap; the robot locks the cable into place. Train lights flicker once, then surge back on, casting warm reflections across the flooded rails. 10-12s: the camera rises and pulls back as the restored monorail crosses a high bridge through rain and fog, revealing the stacked mountain city; end with the robot silhouetted against the restored lights and a distant thunder flash. Sound: close rain impacts, wet rail vibration, low electrical hum, sharp cable-lock click, one controlled power surge, distant thunder. No dialogue, no on-screen text, no logos, no anime style, no distorted train geometry, no extra robots, no extra limbs, no jump cuts.
Three things carry this prompt:
- "Hard visual anchor" plus an explicit subject checklist. Naming the robot, the city, the monorail geometry, the rail material, and the light color before any action starts forces the model to lock those traits in from frame zero, instead of letting them drift once the camera starts moving.
- Timestamped beats (0-3s, 3-7s, 7-10s, 10-12s). Each beat describes one camera position and one physical action — accelerate, track, arc-and-lock, pull back. Seedance 2.0 Ext follows timed beats far more reliably than a single unsegmented action description.
- A negative list that targets this model's actual failure modes, not generic ones: distorted train geometry and extra limbs are the specific artifacts that show up when a robot-on-a-moving-vehicle shot goes wrong, so naming them directly suppresses them.
Technique 2: First-last-frame interpolation
This clip tells almost the same story, but it locks down the ending composition too — useful when the final frame needs to match something specific, like a transition into the next shot or a precise hero composition.
Create one continuous 12-second vertical cinematic shot strictly interpolating between the supplied first frame and supplied final frame. The first image is the exact opening state; the second image is the exact final state. Preserve the same orange-and-white monorail, original orange maintenance robot, repaired orange roof cable, rainy mountain city, wet rail materials, amber lights and blue fog throughout. 0-3s: begin at the first-frame rail-height viewpoint and accelerate forward along the wet rails toward the train roof. 3-7s: track beside the roof as the robot braces against wind, pulls the heavy orange cable to the damaged roof connector; show natural cable tension, metal friction, rain and water spray. 7-9s: brief blue electrical arc, cable locks in, train lights flicker then stabilize. 9-12s: camera moves smoothly behind and above the train until it exactly reaches the supplied final-frame composition: full rear of train moving away, two parallel wet steel rails visibly continuous from foreground beneath the train across the railway viaduct and converging beyond it. Never show the train front in the final section. Keep the railway geometry physically coherent and track continuity visible at every moment. Sound: rain impacts, wet rail vibration, low electrical hum, cable-lock click, controlled power surge, distant thunder. No dialogue, no text, no logos, no anime, no road bridge, no trackless platform, no missing rails, no extra robots, no extra limbs, no jump cuts.
What changes from technique 1:
- "Strictly interpolating between the supplied first frame and supplied final frame" is stated before anything else. This is the instruction that tells the model it has a fixed destination, not just a fixed start — the rest of the prompt describes the path connecting them.
- The ending is over-specified on purpose: "full rear of train moving away," "two parallel wet steel rails visibly continuous," "never show the train front in the final section." When you're pinning an end frame, vague wording about the final composition is exactly where interpolation drifts — the fix is to describe the last 2-3 seconds in as much geometric detail as the opening beat.
- New negative terms specific to the ending shot — "no road bridge, no trackless platform, no missing rails" — because a camera move toward a rear three-quarter view is exactly where this model is prone to losing rail continuity.
Which technique should you use?
Reach for first-frame anchoring when continuity with an existing image matters more than a precise ending — product shots that need to "come alive" from a static photo, or any scene where you want the model to improvise camera work and timing.
Reach for first-last-frame interpolation when the final composition is load-bearing — the frame needs to match a cut point, loop back to a loop-start frame, or land on a specific angle for a thumbnail or the next shot in a sequence. The tradeoff is less improvisational freedom in exchange for a guaranteed landing spot.
Reproducing both techniques with /v1/tasks
Both techniques use the identical model id, seedance-2.0@ext — not the bare seedance-2.0 id, which is a separate route without image-anchoring support. The only difference in the request body is whether you supply last_frame_url.
First-frame anchoring (first_frame_url only):
{
"model": "seedance-2.0@ext",
"input": {
"prompt": "Use the supplied image as the exact first frame and hard visual anchor... (full prompt above)",
"first_frame_url": "https://your-storage.example.com/anchor-frame.jpg",
"aspect_ratio": "9:16",
"duration": 12,
"resolution": "1080p",
"generate_audio": true
}
}
First-last-frame interpolation (first_frame_url + last_frame_url):
{
"model": "seedance-2.0@ext",
"input": {
"prompt": "Create one continuous 12-second vertical cinematic shot strictly interpolating... (full prompt above)",
"first_frame_url": "https://your-storage.example.com/first-frame.jpg",
"last_frame_url": "https://your-storage.example.com/last-frame.jpg",
"aspect_ratio": "9:16",
"duration": 12,
"resolution": "1080p",
"generate_audio": true
}
}
POST either body to /v1/tasks, then poll GET /v1/tasks/{taskId} until status is success and download output[0].url right away — these links expire.
Pricing on the @ext route is driven only by resolution and duration, confirmed against the live pricing endpoint:
| Resolution | Price per second | 12-second clip |
|---|---|---|
| 720p | $0.149 | $1.79 |
| 1080p | $0.373 | $4.48 |
| 4K | $0.467 | $5.60 |
Anchoring one frame or two costs the same — you're not charged extra for last_frame_url.
FAQ
Does seedance-2.0@ext support image-to-video, or is it text-only?
It supports both. first_frame_url alone gives you image-to-video with free camera movement; adding last_frame_url gives you first-last-frame interpolation between two fixed images.
Can I use last_frame_url without first_frame_url?
No — the API requires first_frame_url whenever last_frame_url is present. You can't anchor only the ending.
Why use seedance-2.0@ext instead of the base seedance-2.0 model id?
The @ext route is what carries image-anchoring support (first_frame_url/last_frame_url) and reference-to-video inputs, and it has its own pricing tiers. The base route is a separate, more limited line.
What resolution should I pick for a narrative shot like this? 1080p is the practical default for anything you'll publish — 720p is noticeably softer on fine detail like the cable and rain texture, and 4K roughly triples the 720p cost for a narrative clip where viewers rarely sit at native resolution.
Try it yourself
Both clips above were produced through hiapi's Seedance 2.0 Ext model page, where you can see the full parameter schema and test prompts directly in the playground. If you want more copy-paste starting points for this model, the Seedance 2.0 Ext prompt recipes collects several more, and this one-shot POV flythrough breakdown walks through a third camera-control technique on the same model.









