Choose a model, enter your prompt, and see the result.
HiAPI Blog
HiAPI
Generate it with HiAPI
Most Seedance 2.0 output defaults to looking like a commercial: steady gimbal movement, even lighting, a camera that drifts like it's on a dolly. That "cinematic" look is exactly what gives AI video away. The clip above was built to do the opposite — it's prompted to look like a real clip your sister shot on her phone while you were making tea, complete with focus hunting, a shaky hand, and an abrupt cut at the end.
The technique isn't a special mode or parameter. It's a prompt that describes camera flaws as deliberately as most prompts describe camera quality — and backs that up with an explicit list of the commercial-grade polish to avoid.
This is the complete prompt used to generate the clip above, copy-pasteable as-is:
A young man in his early twenties stands at a cluttered kitchen counter in a small Shanghai apartment, early evening light coming through a fogged window. He's pouring hot water from a kettle into a mug, steam rising, half-glancing toward his sister who is filming him on her phone from across the counter — he knows she's filming but is trying to ignore it, giving an awkward almost-smile.
Camera: handheld smartphone, vertical 9:16. Visible handheld tilt and tremor throughout, not stabilized like a gimbal. Autofocus hunts and drifts once mid-clip — from his eyes to the fogged window behind him — before snapping back. Auto-exposure visibly steps/jumps once as the steam catches the light. Around 3 seconds in, as he dips the spoon quickly into the mug, rolling-shutter warp skews the spoon's motion. Natural phone stabilization edge-warp is visible at the frame boundaries when he shifts weight.
Audio: native phone mic audio — kitchen ambience, kettle sound, faint hallway noise, his sister's muffled giggle off-mic, no added score.
Ending: abrupt cut to black, no fade, like the sister just stopped recording mid-moment.
Explicitly avoid: no tripod or gimbal stabilization, no professional three-point lighting, no commercial/advertising grade polish, no beauty filter or skin smoothing, no centered rule-of-thirds composition, no floating/drone-smooth camera movement, no artificial shallow depth-of-field or bokeh — this must read as real, unplanned phone footage, not a staged ad.
Generic prompts ask for "handheld camera" and get a gentle, controlled sway — the model's default interpretation of "handheld" still looks professional. This prompt instead calls out four specific, independent camera artifacts: tilt/tremor that never settles, a mid-clip autofocus hunt from the subject's eyes to the fogged window, a visible auto-exposure step as the steam catches the light, and rolling-shutter warp on the one fast motion in the frame (the spoon dip). Each of these is a real signature of actual phone hardware fighting to track a moving, imperfectly lit scene — a single blanket instruction like "shaky handheld" produces none of them.
The subject knows he's being filmed and is visibly trying to ignore it — not posing, not performing. That half-aware, half-irritated energy is what separates "caught on camera" from "shot for camera," and it's worth writing into the prompt explicitly rather than leaving it to the model's default of a subject looking confidently at the lens.
The clip cuts to black abruptly, mid-moment, as if the person filming simply stopped. AI video almost always wants to resolve — a settle, a fade, a final beat. Prompting an unresolved ending is a small detail that does a disproportionate amount of work in selling the clip as real.
The prompt asks for native phone-mic audio — kitchen ambience, the kettle, a stray laugh — with no added score. Generated video with a music bed under it reads as produced content by default; ambient sync audio reads as captured.
The explicit "avoid" list — no gimbal, no professional lighting, no beauty filter, no centered composition, no floating camera, no artificial shallow depth of field — isn't boilerplate. Each one blocks a specific default the model leans toward when a prompt doesn't rule it out. Positive description gets you most of the way there; the negative list is what stops the model from "helpfully" smoothing the result back toward a commercial look.
Seedance 2.0 runs on hiapi's unified /v1/tasks endpoint. The input schema for this model takes a prompt, an integer duration (4-15 seconds), a resolution (480p, 720p, 1080p, or 4k), and an aspect_ratio (1:1, 4:3, 3:4, 16:9, 9:16, 21:9, or adaptive):
curl -X POST https://api.hiapi.ai/v1/tasks \
-H "Authorization: Bearer $HIAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2.0",
"input": {
"prompt": "A young man in his early twenties stands at a cluttered kitchen counter...",
"duration": 5,
"resolution": "720p",
"aspect_ratio": "9:16"
}
}'
Poll GET /v1/tasks/{taskId} until status is success, then download the clip from output[0].url before it expires. At 720p, Seedance 2.0 bills $0.293 per second — a 5-second clip like this one runs about $1.47. Current rates for every resolution tier are on the pricing page.
Why does my Seedance 2.0 output always look like a commercial? Because the model's defaults lean toward polish — stable movement, even exposure, a resolved ending — unless the prompt actively specifies otherwise. "Cinematic," "professional," and "high quality" push further in that direction, not away from it.
Can Seedance 2.0 actually simulate camera shake and autofocus hunting? Yes — both showed up reliably in this clip from prompt text alone, with no special parameter. Describe the specific artifact (tilt, focus drift, exposure step, rolling shutter) rather than a generic adjective like "shaky."
What aspect ratio and resolution should I use for a phone-style clip?
9:16 for vertical phone framing, and 720p is usually enough — phone footage isn't supposed to look sharp or high-production anyway. Save 1080p/4k for shots where detail is the point.
How long can a single Seedance 2.0 clip be?
4 to 15 seconds per /v1/tasks call. Candid, phone-style moments usually read better short — 5 seconds is enough for one unresolved beat.
Does adding audio help sell the realism? Yes. Native, ambient mic audio (room noise, a stray sound, no music bed) reads as captured footage; a music score under the clip reads as produced content even if the visuals are convincing.
Want to try this prompt style yourself? Start from the Seedance 2.0 model page or the API quickstart to make your first /v1/tasks call — and see the one-shot POV framing variant in Seedance 2.0: One-Shot POV Video Prompt for a different way to fight the "AI look."