gpt-image-2/text-to-image
OnlineText-to-image with accurate text rendering; @pro / @beta / @ext routes available.
gpt-image-2/text-to-image All models — image, video, and audio — are called through the unified POST /v1/tasks endpoint, with the model selected via the model field and parameters in input. See each model page for its input fields.
Current pricing is maintained on the main site: HiAPI Pricing.
Text-to-image, image-to-image, reference editing, and photorealistic generation.
Text-to-image with accurate text rendering; @pro / @beta / @ext routes available.
gpt-image-2/text-to-image Image-to-image from one or more references; @pro / @ext routes available.
gpt-image-2/image-to-image xAI standard text-to-image: fast, low-cost, 13 aspect ratios.
grok-imagine/text-to-image xAI image editing: reference-based repaint, composite, transform.
grok-imagine/image-to-image xAI quality text-to-image: hero-grade detail, 1k/2k tiers.
grok-imagine-quality/text-to-image xAI quality image editing with 1k/2k tiers.
grok-imagine-quality/image-to-image ByteDance Seedream 4.5 text-to-image: 4K and in-image CJK text.
seedream-4.5/text-to-image Seedream 4.5 image editing with up to 14 reference images.
seedream-4.5/image-to-image Seedream 5.0 Lite text-to-image: budget-friendly volume generation.
seedream-5.0-lite/text-to-image Seedream 5.0 Lite instruction-based image editing.
seedream-5.0-lite/image-to-image Fast, prompt-tolerant text-to-image for quick iteration.
nano-banana Balanced tier with optional references and 1K / 2K / 4K output.
nano-banana-2 Lite tier: low-cost, fast generation.
Nano-Banana-2-Lite Premium brand visuals, reference editing, high-resolution output.
nano-banana-pro FLUX.2 Pro text-to-image: photorealism with design control.
flux-2/text-to-image FLUX.2 image editing: background swap, material edit, multi-image composite.
flux-2/image-to-image Photorealistic generation with optional reference guidance.
flux-1.1-pro FLUX schnell tier: sub-second-class generation.
flux-schnell/text-to-image Low-cost model with strong Chinese prompts and CJK text rendering.
qwen-image-2.0 Qwen image Pro tier: richer detail and layout.
qwen-image-2.0-pro Photoreal portraits and photographic scenes.
z-image Alibaba Wan 2.7 text-to-image with commercial-photo finish.
wan2.7-image/text-to-image Strong at typography and logo design.
ideogram-v4 Text-to-video, image-to-video, reference media, and short-form generation.
Google Veo 3.1 text-to-video with a 4K tier.
veo-3.1/text-to-video Veo 3.1 image-to-video driven by a first frame.
veo-3.1/image-to-video Veo 3.1 Fast text-to-video: quicker and cheaper.
veo-3.1-fast/text-to-video Veo 3.1 Fast image-to-video.
veo-3.1-fast/image-to-video Reference images, video, and audio, optional synced audio, up to 4K.
seedance-2.0 Seedance mini tier: budget 480p / 720p clips.
seedance-2.0-mini Seedance fast tier: speed-first generation.
seedance-2.0-fast Next-gen Seedance, opening soon.
seedance-2.5 xAI text-to-video for fast short clips.
grok-imagine/text-to-video xAI image-to-video that animates a still.
grok-imagine/image-to-video xAI 1.5-gen image-to-video with stronger motion and consistency.
grok-imagine-1.5/image-to-video Short text-to-video, 3-15s, 720p / 1080p, with audio.
happyhorse-1.0 HappyHorse 1.1 text-to-video.
happyhorse-1.1/text-to-video HappyHorse 1.1 image-to-video.
happyhorse-1.1/image-to-video HappyHorse 1.1 reference-to-video with @-referenced subjects.
happyhorse-1.1/reference-to-video MiniMax Hailuo 2.3 text-to-video with strong motion physics.
hailuo-2.3/text-to-video Hailuo 2.3 image-to-video driven by a first frame.
hailuo-2.3/image-to-video Hailuo 2.3 Fast image-to-video: the family price floor.
hailuo-2.3-fast/image-to-video Kling 3.0 Turbo text-to-video, per-second 720p / 1080p tiers.
kling-3.0-turbo/text-to-video Kling 3.0 Turbo image-to-video.
kling-3.0-turbo/image-to-video Kling 3.0 Omni text-to-video, up to 4K with audio.
kling-3.0-omni/text-to-video Kling 3.0 Omni image-to-video.
kling-3.0-omni/image-to-video Text-to-video with size, duration, prompt expansion, and shot controls.
wan2.7-video/text-to-video First-frame driven, with optional last frame, audio, or clip inputs.
wan2.7-video/image-to-video Music generation and multi-speaker dialogue synthesis.
Multi-speaker dialogue TTS with emotion tags and voice casting.
elevenlabs/text-to-dialogue MiniMax music generation from a text brief.
minimax-music-1.5 Next-gen music: auto lyrics, custom lyrics, or instrumental.
minimax-music-2.6