Changelog
July 2026
-
New models: Seedream 5.0 Pro Text to Image and Image to Image — HiAPI flagship quality tier. Upgraded photorealism with sharp in-image text rendering (brush calligraphy and signage verified),
qualitybasic/highoutput 1K/2K images billed per image; editing takes 1-10 reference images. -
New model family: MiniMax Hailuo 2.3 — standard
hailuo-2.3/text-to-videoandimage-to-video(6/10s per-video tiers, motion-physics leader), plushailuo-2.3-fastimage-to-video — the price floor of the family. -
New models: Kling 3.0 Turbo Text to Video and Image to Video — the speed tier of the Kling 3.0 family. 3-15 second clips at 720p/1080p with per-second tiered billing and lip-synced quoted dialogue; image-to-video is first-frame driven.
-
New models: Grok Imagine Text to Image and Image to Image — xAI's image generation line with 13 aspect ratios (ultra-wide 2:1/20:9 to extra-tall 9:20) and reference edits with up to 3 images. Both bill one flat rate per image; the grok-imagine-quality family offers hero-grade output tiered by
1k/2kresolution. -
New models: Seedream 4.5 Text to Image and Image to Image — ByteDance's community-favorite quality tier. Photorealistic detail with accurate in-image text rendering (English and Chinese), 8 aspect ratios, 2K/4K at one flat per-image price; editing takes up to 14 reference images with subject and text consistency.
-
New model: FLUX.2 Image to Image — pro-tier image editing with up to 8 reference images per request. Multi-image composition, background replacement, style and material edits with strong subject and logo consistency;
aspect_ratio: autofollows the first input, 1K/2K output billed per image. -
New model: MiniMax Music 2.6 — full-song music generation with natural vocals in English or Chinese. Lyrics up to 3,500 characters with 14 structure tags, instrumental mode via
is_instrumental, and auto-generated lyrics whenlyricsis left empty. Billed per song. -
POST /v1/tasksnow accepts an optional top-levelrouteparameter for models with multiple routes — passing the bare model name withroute: "pro"is the preferred spelling of the legacy@prosuffix (the old suffix keeps working). An unknown route returns a 400 listing the available routes, and the task detail echoesrouteplus the resolved full model name. See Model routes on the Create Task page. -
New models: five additions across image, video and audio — Nano Banana 2 Lite (entry-tier 1K image generation), Qwen Image 2.0 Pro (pro-grade text rendering and photorealism), Ideogram V4 (best-in-class typography with TURBO/BALANCED/QUALITY tiers), Grok Imagine 1.5 Image to Video preview (upgraded motion, 1–15s), and MiniMax Music 1.5 (full songs up to 4 minutes with vocals, HiAPI's first music model).
-
New models: Seedream 5.0 Lite Text to Image and Image to Image — ByteDance's reasoning-capable image model with real-time web knowledge and accurate multilingual text rendering. Editing supports up to 14 reference images and in-image text rewriting. 2K/4K output at the same per-image price.
-
POST /v1/tasksnow accepts an optionalIdempotency-Keyheader — retries with the same key under the same account create the task only once, and replays return the original taskId, preventing duplicate tasks and duplicate charges from timeout retries. See Idempotency key in the Unified Async API intro. -
New models: Veo 3.1 Text to Video and Image to Video — Google's flagship video tier with cinematic quality, native audio, up to 4K, and 4/6/8-second clips. Also added Veo 3.1 Fast Image to Video, the cost-efficient way to animate stills.
-
New model: ElevenLabs Text to Dialogue v3 — HiAPI's first audio model. Multi-speaker dialogue speech with a voice per line (67 presets), 70+ languages, tunable stability, billed per character.
-
New models: Grok Imagine Text to Video and Image to Video — xAI Grok Imagine video generation with selectable motion mode, aspect ratio, 6–30s duration, and 480p/720p resolution.
-
seedance-2-0now supports4kresolution — pick480p,720p,1080p, or4kfor ultra HD output. -
New model: Seedance 2.0 Fast — ByteDance's high-speed video model for text-to-video, image-to-video (first/last frame), and multimodal reference-to-video (image/video/audio), with native audio. 480p/720p, 4–15s, billed per second.
June 2026
-
New model: Seedance 2.0 Mini — ByteDance's cost-efficient video model for both text-to-video and image-to-video, with native audio, first/last-frame control, and image/video/audio multimodal references. 480p/720p, 4–15s, billed per second.
-
Output Storage is live — keep the images and videos you create with
POST /v1/taskspast the default 7 days. Pick a storage tier when you create a task, or promote existing outputs to long-term storage, then list and delete them from the API or your dashboard. Results are still fetched the same way, viaGET /v1/tasks/:idor acallback.url.
May 2026
-
New model: HappyHorse 1.0 text-to-video — 720p/1080p, 3–15s, multiple aspect ratios.
-
Agent Skills — call a single model straight from Codex, Claude Code, OpenClaw and other agents with one install.
April 2026
-
gpt-image-2added 1K / 2K / 4K resolution tiers. -
Remote MCP server — connect any MCP client and generate with image and video tools over a single endpoint.
-
New models: Qwen Image 2.0 and Seedance 2.0.