AI Models Market: Text, Image, Video & Audio APIs
The HiAPI models market brings together text LLMs, image, video, music, and speech models. Available models include one-key access, usage-based pricing, an online Playground, and request examples; upcoming models include capability and launch previews. Filter by provider or task type.
- DeepSeek V4 Flash API — text generation. DeepSeek V4 Flash is an open-weight language model for high-throughput generation, reasoning, coding, and agent workflows, with a 1M-token context window, thinking and non-thinking modes, JSON output, and tool calling.
- ElevenLabs Text to Dialogue API — music generation, 250 Credits. ElevenLabs Eleven v3 Text-to-Dialogue is the most expressive multi-speaker speech model, turning scripted dialogue into natural, emotionally-rich audio across 70+ languages. Assign a different voice per line, control stability, and generate multi-character conversations in one call.
- FLUX 1.1 Pro API — image generation, 100 Credits. Black Forest Labs' most advanced image generation model with exceptional photorealism and prompt adherence
- flux-2/image-to-image API — image generation, 70 Credits. FLUX.2 pro-tier image editing: combine up to 8 reference images with instruction-precise edits, consistent subjects and logos, 1K/2K output and auto aspect matching to your input.
- FLUX.2 [klein] 4B Image to Image API — image generation, 14 Credits. Compact, fast single-reference FLUX.2 editing with natural-language changes and output up to 4 MP.
- FLUX.2 [klein] 4B Text to Image API — image generation, 2 Credits. A compact and fast FLUX.2 text-to-image model for cost-efficient production, with output from 0.25 to 4 megapixels.
- FLUX.2 [klein] 9B Image to Image API — image generation, 62 Credits. Fast four-step single-reference FLUX.2 editing with stronger detail preservation, text rendering, and commercial polish.
- FLUX.2 [klein] 9B Text to Image API — image generation, 17 Credits. Fast four-step FLUX.2 text-to-image generation with stronger realism, readable typography, and refined detail.
- FLUX.2 Pro API — image generation, 74 Credits. Black Forest Labs FLUX.2 Pro: high-fidelity text-to-image with strong prompt adherence and crisp in-image text.
- flux-3 API — video generation, coming soon. FLUX 3 is Black Forest Labs' next-generation model family. FLUX.3 Video is currently in Early Access, with officially announced support for up to 20-second native-audio video, text-to-video, image/reference-guided video, video-to-video, video-audio continuation, keyframe control, multilingual dialogue, and multi-shot chaining. Access opens after the production API contract and pricing are published.
- FLUX.1 Schnell API — image generation, 10 Credits. FLUX.1 Schnell, Black Forest Labs' open-source ultra-fast text-to-image model: 1-4 inference steps with second-level output, Apache 2.0 commercial license, and HiAPI's lowest per-image price tier — built for high-volume generation and rapid iteration.
- GPT Image 2 Image-to-Image API — image generation, 60 Credits. OpenAI image-to-image model for fast, high-quality image generation and editing with flexible aspect ratios and resolutions.
- GPT Image 2 Multi-ratio 4K I2I API — image generation, 14 Credits. GPT Image 2 image-to-image high-resolution multi-ratio line: blend and edit with up to 16 reference images, across 16 aspect ratios x 1K/2K/4K x low/medium/high quality tiers. Precise control of output framing and finish for editing, style transfer and multi-image fusion.
- GPT Image 2 API — image generation, 60 Credits. OpenAI's next-generation image model with richer detail and more natural color, called via the multimodal Chat Completions API.
- GPT Image 2 Beta API — image generation, 40 Credits. GPT Image 2 Beta is the preview of OpenAI's next-generation image model. It works with the standard OpenAI Images API format and produces high-quality images, ideal for creative design, posters, illustrations, product concepts, and social media assets.
- GPT Image 2 Multi-ratio 4K API — image generation, 14 Credits. GPT Image 2 high-resolution multi-ratio line: freely combine 16 aspect ratios (incl. 5:4, 4:5, 2:1, 21:9) x 1K/2K/4K resolutions x low/medium/high quality tiers, up to 3840x2160 output. Built for posters, banners and print assets that demand exact framing and finish.
- Grok Imagine 1.5 Image to Video API — video generation, 22 Credits. xAI Grok Imagine 1.5 image-to-video: animate a still image with upgraded motion quality, 1-15 seconds, 480p/720p, same price as 1.0.
- grok-imagine/image-to-image API — image generation, 70 Credits. xAI Grok Imagine image editing: reference-driven edits and blends (up to 6 images), auto aspect follows your input, flat per-image billing.
- Grok Imagine Image to Video API — video generation, 22 Credits. xAI Grok Imagine image-to-video animates one or more reference images into a cinematic short clip, with an optional motion prompt, selectable aspect ratio, duration (6-30s), and 480p/720p resolution.
- grok-imagine-quality/image-to-image API — image generation, 140 Credits. xAI Grok Imagine quality-tier image editing: stronger consistency and detail preservation (up to 6 references), tiered 1K/2K pricing.
- grok-imagine-quality/text-to-image API — image generation, 140 Credits. xAI Grok Imagine quality-tier text-to-image: richer detail and aesthetics, tiered per-image pricing by 1K/2K resolution.
- grok-imagine/text-to-image API — image generation, 60 Credits. xAI Grok Imagine text-to-image: fast, low-cost generation with 13 aspect ratios (incl. ultra-wide 2:1/20:9), 1K/2K at one flat per-image price.
- Grok Imagine Text to Video API — video generation, 22 Credits. xAI Grok Imagine text-to-video generates cinematic short clips from a text prompt, with selectable motion mode, aspect ratio, duration (6-30s), and 480p/720p resolution.
- hailuo-2.3-fast/image-to-video API — video generation, 540 Credits. MiniMax Hailuo 2.3 Fast image-to-video: the speed tier at the lowest price, 6s/10s per-video billing.
- hailuo-2.3/image-to-video API — video generation, 800 Credits. MiniMax Hailuo 2.3 image-to-video: first-frame driven with natural motion, 6s/10s clips billed per video.
- hailuo-2.3/text-to-video API — video generation, 800 Credits. MiniMax Hailuo 2.3 text-to-video: the standard tier with excellent motion physics, 6s/10s clips billed per video.
- HappyHorse 1.0 API — video generation, 336 Credits. HappyHorse text-to-video model, supporting 720p / 1080p, 3-15 second durations, and multiple aspect ratios.
- HappyHorse 1.1 Image-to-Video API — video generation, 420 Credits. Alibaba HappyHorse 1.1 image-to-video: drive generation from a first frame while keeping subject and style consistent, native audio, 720p/1080p.
- HappyHorse 1.1 Reference-to-Video API — video generation, 420 Credits. Alibaba HappyHorse 1.1 reference-to-video: generate video from up to 9 reference images while keeping subject, scene, and style consistent, native audio.
- HappyHorse 1.1 Text-to-Video API — video generation, 420 Credits. Alibaba HappyHorse 1.1 text-to-video: cinematic motion, strong prompt adherence, native audio, 720p/1080p.
- Ideogram V4 API — image generation, 200 Credits. Ideogram V4: the typography specialist — best-in-class in-image text accuracy for posters, logos and layouts, with three rendering speed tiers.
- Kling 3.0 Omni Image-to-Video API — video generation, 258 Credits. Kuaishou Kling 3.0 Omni image-to-video: drive generation from first/last frames, native synced audio, up to 4K.
- Kling 3.0 Omni Text-to-Video API — video generation, 258 Credits. Kuaishou Kling 3.0 Omni text-to-video: cinematic motion, multi-shot storytelling, native synced audio, up to 4K.
- kling-3.0-turbo/image-to-video API — video generation, 260 Credits. Kling 3.0 Turbo image-to-video: first-frame driven generation at speed, 3-15s, 720p/1080p, billed per second.
- kling-3.0-turbo/text-to-video API — video generation, 260 Credits. Kling 3.0 Turbo text-to-video: the speed tier of the 3.0 family with flexible 3-15s duration and 720p/1080p output — built for volume.
- minimax-h3 API — video generation, 237 Credits. MiniMax H3 is a native 2K multimodal video model for text-to-video, first/last-frame control, and reference-driven generation with images, video, and audio across 4 to 15 seconds.
- MiniMax Music 1.5 API — music generation, 140 Credits. MiniMax music generation: turn a style prompt and lyrics into a complete song up to about 4 minutes, with natural vocals and rich instrumentation, singing in English or Chinese.
- minimax-music-2.6 API — music generation, 420 Credits. MiniMax Music 2.6: generate complete songs with natural vocals and rich instrumentation from a style prompt and optional lyrics. Supports instrumental mode, automatic lyrics generation, and song-structure tags, singing in English or Chinese.
- Nano Banana API — image generation, 100 Credits. Google's powerful image generation model with stunning quality and fast generation times
- Nano Banana 2 API — image generation, 102 Credits. Nano Banana 2 — the Google Gemini 3.1 Flash Image model. Built for developers, it pairs lightning speed with Pro-grade quality: precise text rendering, strong character consistency, and up to 4K output. The best balance of speed, quality, and price for large-scale image generation and editing workflows. Commercial license supported.
- Nano Banana 2 Lite API — image generation, 66 Credits. Google Nano Banana 2 Lite: low-latency, ultra low-cost 1K image generation with optional reference images (up to 10) for editing and remixing.
- Nano Banana Pro API — image generation, 340 Credits. Nano Banana Pro — the Gemini 3 Pro Image model from Google DeepMind and the flagship of the Nano Banana series. Delivers top-tier output: sharper 2K imagery, smart upscaling, advanced text rendering, and outstanding character consistency. Built for high-end creative work, brand asset generation, and API-driven production workflows. Commercial license supported.
- qwen-audio-3.0-tts-flash API — music generation, coming soon. Qwen-Audio 3.0 low-latency text-to-speech tier for AI assistants, voice agents, and responsive speech workflows, billed by Alibaba's effective character count.
- qwen-audio-3.0-tts-plus API — music generation, coming soon. Qwen-Audio 3.0 quality text-to-speech tier for voiceover, video narration, and spoken content, billed by Alibaba's effective character count.
- Qwen Image 2.0 API — image generation, 50 Credits. Alibaba Qwen Image 2.0 - cost-effective image generation with excellent Chinese text rendering, multi-style output, up to 2K resolution.
- Qwen Image 2.0 Pro API — image generation, 214 Credits. Alibaba Qwen Image 2.0 Pro: the professional tier with stronger in-image text rendering, finer photorealistic detail, and better semantic adherence, up to 2048x2048.
- Seedance 2.0 API — video generation, 272 Credits. ByteDance Seedance 2.0 - ByteDance's latest video generation model with cinematic quality, exceptional motion, and native audio.
- Seedance 2.0 Ext API — video generation, 298 Credits. Seedance 2.0 extended route for cinematic video generation up to 4K, with native audio plus image and audio references.
- Seedance 2.0 Fast API — video generation, 471 Credits. ByteDance Seedance 2.0 Fast: high-speed video generation with native audio. Text-to-video, image-to-video (first/last frame), and multimodal reference-to-video (image/video/audio). 480p/720p, 4-15s.
- Seedance 2.0 Mini API — video generation, 294 Credits. Seedance 2.0 Mini by ByteDance is a cost-efficient video generation model supporting both text-to-video and image-to-video, with native synced audio, first/last-frame control, and image/video/audio multimodal references. Up to 720P, flexible 4–15s clips.
- Seedance 2.5 API — video generation, coming soon. Seedance 2.5 is ByteDance's next-generation video model supporting up to 30-second single-shot generation, native 4K resolution, up to 50 reference inputs, and 3D blockout preview. Coming soon to HiAPI — stay tuned.
- seedream-4.5/image-to-image API — image generation, 90 Credits. ByteDance Seedream 4.5 image editing: unified generation-editing architecture with up to 14 reference images for edits and composites, consistent subjects, 2K/4K output.
- seedream-4.5/text-to-image API — image generation, 90 Credits. ByteDance Seedream 4.5 text-to-image: the community-favorite quality tier with solid photorealism and in-image text rendering, 2K/4K output and 8 aspect ratios.
- Seedream 5.0 Lite Image to Image API — image generation, 70 Credits. ByteDance Seedream 5.0 Lite image editing: instruction-based edits with up to 14 reference images, following edit instructions precisely while keeping non-edited areas consistent.
- Seedream 5.0 Lite Text to Image API — image generation, 70 Credits. ByteDance Seedream 5.0 Lite text-to-image: reasoning-guided generation with real-time web knowledge, precise instruction following, and strong multilingual text rendering.
- Seedream 5.0 Pro Image to Image API — image generation, 120 Credits. ByteDance Seedream 5.0 Pro image editing: up to 10 reference images for character-, product- and style-consistent edits and composites, 1K/2K output.
- Seedream 5.0 Pro Text to Image API — image generation, 100 Credits. ByteDance Seedream 5.0 Pro flagship text-to-image: upgraded fidelity and in-image text rendering with sharp detail, 1K/2K output and 8 aspect ratios.
- Veo 3.1 Fast Image to Video API — video generation, 500 Credits. Google Veo 3.1 Fast image-to-video: fast, cost-efficient image animation with native audio, up to 4K, 4/6/8-second clips.
- Veo 3.1 Fast Text to Video API — video generation, 500 Credits. Google Veo 3.1 Fast: high-speed text-to-video with native audio, up to 4K, supports 4/6/8-second clips.
- Veo 3.1 Image to Video API — video generation, 1,140 Credits. Google Veo 3.1 image-to-video: animate a still image into a cinematic clip with native audio, up to 4K, 4/6/8 seconds.
- Veo 3.1 Text to Video API — video generation, 1,140 Credits. Google Veo 3.1: flagship text-to-video with native audio, cinematic realism, up to 4K, 4/6/8-second clips.
- Wan 2.7 Image Pro API — image generation, 160 Credits. Tongyi Wanxiang Wan 2.7 Image Pro: text-to-image with up to 4K output, strong prompt adherence, and support for both Chinese and English prompts.
- Wan 2.7 Image-to-Video API — video generation, 334 Credits. Animate any still image into a high-quality video with natural motion, supporting first-frame, first+last-frame, and video continuation modes
- Wan 2.7 Text-to-Video API — video generation, 334 Credits. Alibaba's latest video generation model with cinematic quality, native audio support, and up to 1080P 15-second output
- Z-Image API — image generation, 16 Credits. Tongyi Z-Image: efficient photorealistic text-to-image, fast Turbo generation, accurate in-image text rendering in both Chinese and English.
Chat Completions APIs
- DeepSeek V4 Flash API (DeepSeek)
Text to Speech APIs
- ElevenLabs Text to Dialogue API (ElevenLabs)
- qwen-audio-3.0-tts-flash API (Alibaba)
- qwen-audio-3.0-tts-plus API (Alibaba)
Text to Image APIs
- FLUX 1.1 Pro API (Black Forest Labs)
- FLUX.2 [klein] 4B Text to Image API (Black Forest Labs)
- FLUX.2 [klein] 9B Text to Image API (Black Forest Labs)
- FLUX.2 Pro API (Black Forest Labs)
- FLUX.1 Schnell API (Black Forest Labs)
- GPT Image 2 API (OpenAI)
- GPT Image 2 Beta API (OpenAI)
- GPT Image 2 Multi-ratio 4K API (OpenAI)
- grok-imagine-quality/text-to-image API (Grok)
- grok-imagine/text-to-image API (Grok)
- Ideogram V4 API (Ideogram)
- Nano Banana API (Google)
- Nano Banana 2 API (Google)
- Nano Banana 2 Lite API (Google)
- Nano Banana Pro API (Google)
- Qwen Image 2.0 API (Alibaba)
- Qwen Image 2.0 Pro API (Alibaba)
- seedream-4.5/text-to-image API (ByteDance)
- Seedream 5.0 Lite Text to Image API (ByteDance)
- Seedream 5.0 Pro Text to Image API (ByteDance)
- Wan 2.7 Image Pro API (Alibaba)
- Z-Image API (Alibaba)
Image to Image & Editing APIs
- FLUX 1.1 Pro API (Black Forest Labs)
- flux-2/image-to-image API (Black Forest Labs)
- FLUX.2 [klein] 4B Image to Image API (Black Forest Labs)
- FLUX.2 [klein] 9B Image to Image API (Black Forest Labs)
- GPT Image 2 Image-to-Image API (OpenAI)
- GPT Image 2 Multi-ratio 4K I2I API (OpenAI)
- grok-imagine/image-to-image API (Grok)
- grok-imagine-quality/image-to-image API (Grok)
- Nano Banana 2 API (Google)
- Nano Banana 2 Lite API (Google)
- Nano Banana Pro API (Google)
- seedream-4.5/image-to-image API (ByteDance)
- Seedream 5.0 Lite Image to Image API (ByteDance)
- Seedream 5.0 Pro Image to Image API (ByteDance)
Text to Video APIs
- flux-3 API (Black Forest Labs)
- Grok Imagine Text to Video API (Grok)
- hailuo-2.3/text-to-video API (MiniMax)
- HappyHorse 1.0 API (Alibaba)
- HappyHorse 1.1 Text-to-Video API (Alibaba)
- Kling 3.0 Omni Text-to-Video API (Kuaishou)
- kling-3.0-turbo/text-to-video API (Kuaishou)
- minimax-h3 API (MiniMax)
- Seedance 2.0 API (ByteDance)
- Seedance 2.0 Ext API (ByteDance)
- Seedance 2.0 Fast API (ByteDance)
- Seedance 2.0 Mini API (ByteDance)
- Seedance 2.5 API (ByteDance)
- Veo 3.1 Fast Text to Video API (Google)
- Veo 3.1 Text to Video API (Google)
- Wan 2.7 Text-to-Video API (Alibaba)
Image to Video APIs
- flux-3 API (Black Forest Labs)
- Grok Imagine 1.5 Image to Video API (Grok)
- Grok Imagine Image to Video API (Grok)
- hailuo-2.3-fast/image-to-video API (MiniMax)
- hailuo-2.3/image-to-video API (MiniMax)
- HappyHorse 1.1 Image-to-Video API (Alibaba)
- HappyHorse 1.1 Reference-to-Video API (Alibaba)
- Kling 3.0 Omni Image-to-Video API (Kuaishou)
- kling-3.0-turbo/image-to-video API (Kuaishou)
- minimax-h3 API (MiniMax)
- Seedance 2.0 API (ByteDance)
- Seedance 2.0 Ext API (ByteDance)
- Seedance 2.0 Fast API (ByteDance)
- Seedance 2.0 Mini API (ByteDance)
- Seedance 2.5 API (ByteDance)
- Veo 3.1 Fast Image to Video API (Google)
- Veo 3.1 Image to Video API (Google)
- Wan 2.7 Image-to-Video API (Alibaba)
Reference to Video APIs
- flux-3 API (Black Forest Labs)
- minimax-h3 API (MiniMax)
- Seedance 2.0 API (ByteDance)
- Seedance 2.0 Fast API (ByteDance)
- Seedance 2.0 Mini API (ByteDance)
- Seedance 2.5 API (ByteDance)
- Wan 2.7 Image-to-Video API (Alibaba)
Text to Music APIs
- MiniMax Music 1.5 API (MiniMax)
- minimax-music-2.6 API (MiniMax)
DeepSeek Model APIs
- DeepSeek V4 Flash API — text generation
ElevenLabs Model APIs
- ElevenLabs Text to Dialogue API — music generation
Black Forest Labs Model APIs
- FLUX 1.1 Pro API — image generation
- flux-2/image-to-image API — image generation
- FLUX.2 [klein] 4B Image to Image API — image generation
- FLUX.2 [klein] 4B Text to Image API — image generation
- FLUX.2 [klein] 9B Image to Image API — image generation
- FLUX.2 [klein] 9B Text to Image API — image generation
- FLUX.2 Pro API — image generation
- flux-3 API — video generation
- FLUX.1 Schnell API — image generation
OpenAI Model APIs
- GPT Image 2 Image-to-Image API — image generation
- GPT Image 2 Multi-ratio 4K I2I API — image generation
- GPT Image 2 API — image generation
- GPT Image 2 Beta API — image generation
- GPT Image 2 Multi-ratio 4K API — image generation
Grok Model APIs
- Grok Imagine 1.5 Image to Video API — video generation
- grok-imagine/image-to-image API — image generation
- Grok Imagine Image to Video API — video generation
- grok-imagine-quality/image-to-image API — image generation
- grok-imagine-quality/text-to-image API — image generation
- grok-imagine/text-to-image API — image generation
- Grok Imagine Text to Video API — video generation
MiniMax Model APIs
- hailuo-2.3-fast/image-to-video API — video generation
- hailuo-2.3/image-to-video API — video generation
- hailuo-2.3/text-to-video API — video generation
- minimax-h3 API — video generation
- MiniMax Music 1.5 API — music generation
- minimax-music-2.6 API — music generation
Alibaba Model APIs
- HappyHorse 1.0 API — video generation
- HappyHorse 1.1 Image-to-Video API — video generation
- HappyHorse 1.1 Reference-to-Video API — video generation
- HappyHorse 1.1 Text-to-Video API — video generation
- qwen-audio-3.0-tts-flash API — music generation
- qwen-audio-3.0-tts-plus API — music generation
- Qwen Image 2.0 API — image generation
- Qwen Image 2.0 Pro API — image generation
- Wan 2.7 Image Pro API — image generation
- Wan 2.7 Image-to-Video API — video generation
- Wan 2.7 Text-to-Video API — video generation
- Z-Image API — image generation
Ideogram Model APIs
- Ideogram V4 API — image generation
Kuaishou Model APIs
- Kling 3.0 Omni Image-to-Video API — video generation
- Kling 3.0 Omni Text-to-Video API — video generation
- kling-3.0-turbo/image-to-video API — video generation
- kling-3.0-turbo/text-to-video API — video generation
Google Model APIs
- Nano Banana API — image generation
- Nano Banana 2 API — image generation
- Nano Banana 2 Lite API — image generation
- Nano Banana Pro API — image generation
- Veo 3.1 Fast Image to Video API — video generation
- Veo 3.1 Fast Text to Video API — video generation
- Veo 3.1 Image to Video API — video generation
- Veo 3.1 Text to Video API — video generation
ByteDance Model APIs
- Seedance 2.0 API — video generation
- Seedance 2.0 Ext API — video generation
- Seedance 2.0 Fast API — video generation
- Seedance 2.0 Mini API — video generation
- Seedance 2.5 API — video generation
- seedream-4.5/image-to-image API — image generation
- seedream-4.5/text-to-image API — image generation
- Seedream 5.0 Lite Image to Image API — image generation
- Seedream 5.0 Lite Text to Image API — image generation
- Seedream 5.0 Pro Image to Image API — image generation
- Seedream 5.0 Pro Text to Image API — image generation