Most "video prompt" advice stops at describing a scene. That works for a single static shot, but it falls apart the moment you ask a model to hold a continuous 15-second take: the subject drifts, the lighting shifts halfway through, and the ending looks nothing like the opening. The prompt below is a real one — pulled from HiAPI's own asset library, not written for this article — that keeps a single glass product steady through a slow camera arc, a three-layer internal transformation, and a calm final hold, without a single cut.
We're reusing the original render here, not generating a new one, so what you see is exactly what the prompt produced on the first pass.
The Full Prompt
This is the complete, unedited prompt — copy it as a starting template for your own product shot:
One continuous fifteen-second premium technology product film. Hold the exact opening composition for the first second. From second one through second ten, the camera performs a very slow, restrained 12-degree arc to the right around the centered translucent glass monolith while maintaining nearly constant subject scale. During the arc, the dense inner architecture separates with precise mechanical elegance into three ultra-thin luminous layers representing image, video, and audio; the layers glide in orderly parallel paths, then converge into one unified API core. The faceted outer glass shell progressively consolidates with controlled physical continuity into the simpler rounded glass portal shown in the closing keyframe. Cyan-teal and warm amber reflections travel subtly across the glass while the black studio and pedestal remain stable. From second ten to twelve, the camera decelerates smoothly into the closing viewpoint. From second twelve to fifteen, hold the final product frame calm and nearly motionless for a brand title overlay. The sequence stays centered, elegant, quiet, photorealistic, and stabilized; the clean frame contains only the glass product and pedestal, with coherent material, shadows, perspective, and lighting throughout.
Why It Works, Beat by Beat
The prompt reads like a shot list with a stopwatch, not a mood board. That's the core technique: every instruction is anchored to a second range, so the model has no ambiguity about when something should happen.
Second 0–1: hold before you move
"Hold the exact opening composition for the first second" gives the model a clean anchor frame before any motion starts. Skipping this beat is the most common reason AI video opens look shaky — the camera and subject are still negotiating position in the first frame instead of starting from one that's already settled.
Seconds 1–10: one move, described once
The 12-degree arc is deliberately small and covers the longest span of the clip. "Maintaining nearly constant subject scale" is the line doing the real work — it tells the model not to punch in or drift closer as it rotates, which is the usual failure mode for arc shots (the subject creeps larger frame by frame until the composition breaks by second eight).
The transformation is layered, not simultaneous
Inside that same ten-second window, three separate visual ideas happen in parallel: the internal layers (image/video/audio) separate and re-converge, and the outer shell simplifies into the closing keyframe's shape. Naming each layer explicitly — rather than saying "the product transforms" — is what keeps three unrelated visual changes from blurring into one messy morph.
Seconds 10–12: decelerate before you land
A hard stop on a moving camera reads as a mistake, not a choice. Giving the model an explicit two-second deceleration beat before the final composition is what makes the ending feel intentional instead of clipped.
Seconds 12–15: the empty frame is the point
The last three seconds ask for almost no motion at all — "hold the final product frame calm and nearly motionless for a brand title overlay." That's a production instruction, not a creative one: it reserves clean, static screen real estate for whatever text or logo gets added afterward. The clip in the embed above has no baked-in title text — that's intentional. The version with brand typography is composited in post, on top of this same held frame, which is also why you should never ask a video model to render your actual product name or tagline as on-screen text inside the generation itself.
Continuity anchors: the details that stop drift
"Cyan-teal and warm amber reflections," "the black studio and pedestal remain stable," and "coherent material, shadows, perspective, and lighting throughout" aren't decorative flourishes — they're drift prevention. On a continuous 15-second shot, a model will happily reinterpret your lighting palette or pedestal color halfway through if you don't restate what should stay the same, not just what should change.
Reproducing It: An API Example
This shot was generated as image-to-video on seedance-2.0, anchored with a first frame (the opening composition) and a last frame (the closing keyframe) rather than from a text prompt alone — that's what lets the model consolidate the shell shape toward a specific target instead of guessing at one. Here's the equivalent call against /v1/tasks:
curl -X POST https://api.hiapi.ai/v1/tasks \
-H "Authorization: Bearer $HIAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2.0",
"input": {
"prompt": "One continuous fifteen-second premium technology product film. Hold the exact opening composition for the first second. From second one through second ten, the camera performs a very slow, restrained 12-degree arc to the right around the centered subject while maintaining nearly constant subject scale...",
"first_frame_url": "https://your-cdn.example.com/opening-frame.jpg",
"last_frame_url": "https://your-cdn.example.com/closing-frame.jpg",
"duration": 15,
"resolution": "1080p",
"aspect_ratio": "16:9"
}
}'
Poll the returned task until it reaches a terminal state, then pull the finished clip from output[0].url. At 1080p, seedance-2.0 bills $0.729 per second — a 15-second clip like this one costs $10.94 total. If you're prototyping the timing beats before committing to a full-price render, drop to 480p ($0.136/s, about $2.04 for 15 seconds) to check pacing first, then re-render the final pass at 1080p.
FAQ
Can I download this exact clip and use it as my own product video? No — this is HiAPI's internal asset, shown here to demonstrate the prompt technique. Use your own product photography as the first/last frame and adapt the prompt's timing structure to your subject.
Why break the prompt into second-by-second beats instead of one flowing description? Because a 15-second continuous shot has to survive several distinct events (a hold, an arc, an internal transformation, a deceleration, a final hold) without any cuts to reset the model's interpretation. Timestamping each beat is what keeps the physics and composition consistent across all of them instead of the model picking one interpretation and drifting from it.
What if I don't have a first and last frame ready?
seedance-2.0 also runs text-to-video from the prompt alone — drop the first_frame_url/last_frame_url fields and keep duration, resolution, and aspect_ratio. You'll have less control over the exact opening and closing composition, so tighten the prompt's opening-frame description to compensate.
Will the model render my title text directly onto the video? Don't rely on it — video models are far less reliable at rendering clean on-screen typography than image models are. That's exactly why this prompt asks for a near-static held frame at the end instead of asking the model to draw text: composite your title in post over that clean frame.
Try It Yourself
If you want the full parameter list — resolution tiers, reference-video and reference-audio inputs, pricing at every resolution — the seedance-2.0 model page has the live schema and the current pricing breakdown. For a single-shot POV take on the same model with a different camera technique, see our one-shot POV prompt breakdown; for a wider survey of short-form product video use cases beyond this one prompt, see our product short video guide.









