AI Models Market: Text, Image, Video & Audio APIs
The HiAPI models market brings together text LLMs, image, video, music, and speech models. Available models include one-key access, usage-based pricing, an online Playground, and request examples; upcoming models include capability and launch previews. Filter by provider or task type.
- DeepSeek V4 Flash API — text generation. DeepSeek V4 Flash is an open-weight language model for high-throughput generation, reasoning, coding, and agent workflows, with a 1M-token context window, thinking and non-thinking modes, JSON output, and tool calling.
- DeepSeek V4 Pro API — text generation. DeepSeek V4 Pro (0813) targets complex reasoning, coding, and agent workflows. It natively supports the Responses API and remains compatible with Chat Completions.
- ElevenLabs Text to Dialogue API — music generation, 250 Credits. ElevenLabs Eleven v3 Text-to-Dialogue is the most expressive multi-speaker speech model, turning scripted dialogue into natural, emotionally-rich audio across 70+ languages. Assign a different voice per line, control stability, and generate multi-character conversations in one call.
- FLUX 1.1 Pro API — image generation, 100 Credits. Black Forest Labs' most advanced image generation model with exceptional photorealism and prompt adherence
- flux-2/image-to-image API — image generation, 70 Credits. FLUX.2 pro-tier image editing: combine up to 8 reference images with instruction-precise edits, consistent subjects and logos, 1K/2K output and auto aspect matching to your input.
- FLUX.2 [klein] 4B Image to Image API — image generation, 14 Credits. Compact, fast single-reference FLUX.2 editing with natural-language changes and output up to 4 MP.
- FLUX.2 [klein] 4B Text to Image API — image generation, 2 Credits. A compact and fast FLUX.2 text-to-image model for cost-efficient production, with output from 0.25 to 4 megapixels.
- FLUX.2 [klein] 9B Image to Image API — image generation, 62 Credits. Fast four-step single-reference FLUX.2 editing with stronger detail preservation, text rendering, and commercial polish.
- FLUX.2 [klein] 9B Text to Image API — image generation, 17 Credits. Fast four-step FLUX.2 text-to-image generation with stronger realism, readable typography, and refined detail.
- FLUX.2 Pro API — image generation, 74 Credits. Black Forest Labs FLUX.2 Pro: high-fidelity text-to-image with strong prompt adherence and crisp in-image text.
- FLUX.3 Video API — video generation, 500 Credits. Try the FLUX.3 Video API with text prompts, 1 to 10 keyframes, source-video continuation, synchronized audio, Draft previews, and per-second pricing.
- FLUX.1 Schnell API — image generation, 10 Credits. FLUX.1 Schnell, Black Forest Labs' open-source ultra-fast text-to-image model: 1-4 inference steps with second-level output, Apache 2.0 commercial license, and HiAPI's lowest per-image price tier — built for high-volume generation and rapid iteration.
- GPT-5.6 Luna API — text generation. GPT-5.6 Luna targets cost-sensitive, high-throughput workloads such as bulk classification, extraction, rewriting, and lightweight coding. HiAPI currently exposes text, multi-turn, and function tool calling through the streaming Responses API.
- GPT-5.6 Sol API — text generation. GPT-5.6 Sol is the flagship reasoning model in OpenAI's GPT-5.6 family for complex coding, professional analysis, and demanding agent workflows. HiAPI currently exposes text, multi-turn, and function tool calling through the streaming Responses API.
- GPT-5.6 Terra API — text generation. GPT-5.6 Terra balances intelligence, speed, and cost for everyday coding, document work, and general agent workflows. HiAPI currently exposes text, multi-turn, and function tool calling through the streaming Responses API.
- GPT Image 2 API — image generation, 60 Credits. Use the GPT Image 2 text-to-image model on HiAPI for posters, knowledge cards, product visuals, and other generated images. Try it online or integrate the async API.
- GPT Image 2 Image-to-Image API — image generation, 60 Credits. OpenAI image-to-image model for fast, high-quality image generation and editing with flexible aspect ratios and resolutions.
- GPT Image 2 Multi-ratio 4K I2I API — image generation, 14 Credits. GPT Image 2 image-to-image high-resolution multi-ratio line: blend and edit with up to 6 reference images, across 16 aspect ratios x 1K/2K/4K x low/medium/high quality tiers. Precise control of output framing and finish for editing, style transfer and multi-image fusion.
- GPT Image 2 Beta API — image generation, 40 Credits. GPT Image 2 Beta is the preview of OpenAI's next-generation image model. It works with the standard OpenAI Images API format and produces high-quality images, ideal for creative design, posters, illustrations, product concepts, and social media assets.
- GPT Image 2 Multi-ratio 4K API — image generation, 14 Credits. GPT Image 2 high-resolution multi-ratio line: freely combine 16 aspect ratios (incl. 5:4, 4:5, 2:1, 21:9) x 1K/2K/4K resolutions x low/medium/high quality tiers, up to 3840x2160 output. Built for posters, banners and print assets that demand exact framing and finish.
- Grok Imagine 1.5 Image to Video API — video generation, 30 Credits. xAI Grok Imagine 1.5 image-to-video: animate a still image with upgraded motion quality, 1-15 seconds, 480p/720p, same price as 1.0.
- grok-imagine-image-2.0/image-to-image API — image generation, 130 Credits. xAI Grok Imagine Image 2.0 image editing: provide one reference image and describe the edit in natural language, with 1k/2k output.
- grok-imagine-image-2.0/text-to-image API — image generation, 104 Credits. xAI Grok Imagine Image 2.0 text-to-image: generate images with low or medium quality, 1k/2k resolution, and flexible square-to-ultrawide aspect ratios.
- grok-imagine/image-to-image API — image generation, 60 Credits. xAI Grok Imagine image editing: reference-guided edits and multi-image blending with natural-language instructions.
- Grok Imagine Image to Video API — video generation, 30 Credits. xAI Grok Imagine image-to-video animates one or more reference images into a cinematic short clip, with an optional motion prompt, selectable aspect ratio, duration (6-30s), and 480p/720p resolution.
- grok-imagine-quality/image-to-image API — image generation, 160 Credits. xAI Grok Imagine quality-tier image editing: enhanced detail and subject consistency for natural-language edits and multi-image blending.
- grok-imagine-quality/text-to-image API — image generation, 140 Credits. xAI Grok Imagine quality-tier text-to-image: richer detail and aesthetics, tiered per-image pricing by 1K/2K resolution.
- grok-imagine/text-to-image API — image generation, 60 Credits. xAI Grok Imagine text-to-image: fast, low-cost generation with 13 aspect ratios (incl. ultra-wide 2:1/20:9), 1K/2K at one flat per-image price.
- Grok Imagine Text to Video API — video generation, 30 Credits. xAI Grok Imagine text-to-video generates cinematic short clips from a text prompt, with selectable motion mode, aspect ratio, duration (6-30s), and 480p/720p resolution.
- hailuo-2.3-fast/image-to-video API — video generation, 540 Credits. MiniMax Hailuo 2.3 Fast image-to-video: the speed tier at the lowest price, 6s/10s per-video billing.
- hailuo-2.3/image-to-video API — video generation, 800 Credits. MiniMax Hailuo 2.3 image-to-video: first-frame driven with natural motion, 6s/10s clips billed per video.
- hailuo-2.3/text-to-video API — video generation, 800 Credits. MiniMax Hailuo 2.3 text-to-video: the standard tier with excellent motion physics, 6s/10s clips billed per video.
- HappyHorse 1.0 API — video generation, 336 Credits. HappyHorse text-to-video model, supporting 720p / 1080p, 3-15 second durations, and multiple aspect ratios.
- HappyHorse 1.1 Image-to-Video API — video generation, 420 Credits. Alibaba HappyHorse 1.1 image-to-video: drive generation from a first frame while keeping subject and style consistent, native audio, 720p/1080p.
- HappyHorse 1.1 Reference-to-Video API — video generation, 420 Credits. Alibaba HappyHorse 1.1 reference-to-video: generate video from up to 9 reference images while keeping subject, scene, and style consistent, native audio.
- HappyHorse 1.1 Text-to-Video API — video generation, 420 Credits. Alibaba HappyHorse 1.1 text-to-video: cinematic motion, strong prompt adherence, native audio, 720p/1080p.
- Ideogram V4 API — image generation, 200 Credits. Ideogram V4: the typography specialist — best-in-class in-image text accuracy for posters, logos and layouts, with three rendering speed tiers.
- Kling 3.0 Omni Image-to-Video API — video generation, 258 Credits. Kuaishou Kling 3.0 Omni image-to-video: drive generation from first/last frames, native synced audio, up to 4K.
- Kling 3.0 Omni Text-to-Video API — video generation, 258 Credits. Kuaishou Kling 3.0 Omni text-to-video: cinematic motion, multi-shot storytelling, native synced audio, up to 4K.
- Kling 3.0 Turbo Image-to-Video API — video generation, 260 Credits. Kling 3.0 Turbo image-to-video: first-frame driven generation at speed, 3-15s, 720p/1080p, billed per second.
- Kling 3.0 Turbo Text to Video API — video generation, 260 Credits. Try Kling 3.0 Turbo text-to-video in the HiAPI Playground. Compare real 720p results, copy tested prompts, review 3–15 second parameters and pricing, then call POST /v1/tasks.
- minimax-h3 API — video generation, 237 Credits. MiniMax H3 is a native 2K multimodal video model for text-to-video, first/last-frame control, and reference-driven generation with images, video, and audio across 4 to 15 seconds.
- MiniMax Music 1.5 API — music generation, 140 Credits. MiniMax music generation: turn a style prompt and lyrics into a complete song up to about 4 minutes, with natural vocals and rich instrumentation, singing in English or Chinese.
- minimax-music-2.6 API — music generation, 420 Credits. MiniMax Music 2.6: generate complete songs with natural vocals and rich instrumentation from a style prompt and optional lyrics. Supports instrumental mode, automatic lyrics generation, and song-structure tags, singing in English or Chinese.
- minimax-music-3 API — music generation, 5 Credits. MiniMax Music 3 generates complete songs up to five minutes from a music description and lyrics, with detailed control over genre, mood, vocals, instrumentation, and arrangement, output as 44.1 kHz stereo WAV.
- Nano Banana API — image generation, 100 Credits. Google's powerful image generation model with stunning quality and fast generation times
- Nano Banana 2 API — image generation, 102 Credits. Nano Banana 2 — the Google Gemini 3.1 Flash Image model. Built for developers, it pairs lightning speed with Pro-grade quality: precise text rendering, strong character consistency, and up to 4K output. The best balance of speed, quality, and price for large-scale image generation and editing workflows. Commercial license supported.
- Nano Banana 2 Lite API — image generation, 66 Credits. Google Nano Banana 2 Lite: low-latency, ultra low-cost 1K image generation with optional reference images (up to 10) for editing and remixing.
- Nano Banana Pro API — image generation, 340 Credits. Nano Banana Pro — the Gemini 3 Pro Image model from Google DeepMind and the flagship of the Nano Banana series. Delivers top-tier output: sharper 2K imagery, smart upscaling, advanced text rendering, and outstanding character consistency. Built for high-end creative work, brand asset generation, and API-driven production workflows. Commercial license supported.
- Qwen-Audio 3.0 TTS Flash API — music generation, 60 Credits. Qwen-Audio 3.0 low-latency text-to-speech for AI assistants, voice agents, notifications, and responsive speech workflows, billed by effective character count.
- Qwen-Audio 3.0 TTS Plus API — music generation, 80 Credits. Qwen-Audio 3.0 high-quality text-to-speech for voiceovers, video narration, and spoken content, billed by effective character count.
- Qwen Image 2.0 API — image generation, 84 Credits. Alibaba Qwen Image 2.0 - cost-effective image generation with excellent Chinese text rendering, multi-style output, up to 2K resolution.
- Qwen Image 2.0 Pro API — image generation, 214 Credits. Alibaba Qwen Image 2.0 Pro: the professional tier with stronger in-image text rendering, finer photorealistic detail, and better semantic adherence, up to 2048x2048.
- qwen-image-3.0/image-to-image API — image generation, 71 Credits. Qwen Image 3.0 image-to-image uses 1–3 references for redraws, composition preservation, and visual refinement.
- qwen-image-3.0-pro/image-to-image API — image generation, 99 Credits. Qwen Image 3.0 Pro image-to-image uses 1–3 references for high-fidelity redraws, complex material refinement, and detailed layout control.
- qwen-image-3.0-pro/text-to-image API — image generation, 99 Credits. Qwen Image 3.0 Pro text-to-image for high-fidelity final images, complex materials, and refined layouts.
- qwen-image-3.0/text-to-image API — image generation, 71 Credits. Qwen Image 3.0 text-to-image for Chinese text, posters, concepts, and everyday visual creation.
- Seedance 2.0 API — video generation, 272 Credits. ByteDance Seedance 2.0 - ByteDance's latest video generation model with cinematic quality, exceptional motion, and native audio.
- Seedance 2.0 Ext API — video generation, 298 Credits. Seedance 2.0 extended route for cinematic video generation up to 4K, with native audio plus image and audio references.
- Seedance 2.0 Fast API — video generation, 354 Credits. ByteDance Seedance 2.0 Fast: high-speed video generation with native audio. Text-to-video, image-to-video (first/last frame), and multimodal reference-to-video (image/video/audio). 480p/720p, 4-15s.
- Seedance 2.0 Mini API — video generation, 117 Credits. Seedance 2.0 Mini by ByteDance is a cost-efficient video generation model supporting both text-to-video and image-to-video, with native synced audio, first/last-frame control, and image/video/audio multimodal references. Up to 720P, flexible 4–15s clips.
- Seedance 2.5 Image to Video API — video generation, 279 Credits. Try Seedance 2.5 image-to-video with first-frame and first/last-frame control. Compare real 720p results, copy tested prompts, review 4–30 second pricing, and call POST /v1/tasks.
- Seedance 2.5 Reference to Video API — video generation, 242 Credits. Try Seedance 2.5 reference-to-video with video, image, and audio references. Learn @video1 mappings, review 4–30 second pricing, and call POST /v1/tasks.
- Seedance 2.5 Text to Video API — video generation, 462 Credits. ByteDance Seedance 2.5 text-to-video turns a single prompt into one continuous shot of up to 30 seconds at 720p or 1080p, across seven aspect ratios, with natively synchronised audio and prompts in 11 languages.
- Seedream 4.5 Image to Image API — image generation, 90 Credits. ByteDance Seedream 4.5 image editing: unified generation-editing architecture with up to 14 reference images for edits and composites, consistent subjects, 2K/4K output.
- Seedream 4.5 Text to Image API — image generation, 90 Credits. ByteDance Seedream 4.5 text-to-image: the community-favorite quality tier with solid photorealism and in-image text rendering, 2K/4K output and 8 aspect ratios.
- Seedream 5.0 Lite Image to Image API — image generation, 70 Credits. ByteDance Seedream 5.0 Lite image editing: instruction-based edits with up to 14 reference images, following edit instructions precisely while keeping non-edited areas consistent.
- Seedream 5.0 Lite Text to Image API — image generation, 70 Credits. ByteDance Seedream 5.0 Lite text-to-image: reasoning-guided generation with real-time web knowledge, precise instruction following, and strong multilingual text rendering.
- Seedream 5.0 Pro Image to Image API — image generation, 120 Credits. ByteDance Seedream 5.0 Pro image editing: up to 10 reference images for character-, product- and style-consistent edits and composites, 1K/2K output.
- Seedream 5.0 Pro Text to Image API — image generation, 100 Credits. ByteDance Seedream 5.0 Pro flagship text-to-image: upgraded fidelity and in-image text rendering with sharp detail, 1K/2K output and 8 aspect ratios.
- Veo 3.1 Fast Image to Video API — video generation, 500 Credits. Google Veo 3.1 Fast image-to-video: fast, cost-efficient image animation with native audio, up to 4K, 4/6/8-second clips.
- Veo 3.1 Fast Text to Video API — video generation, 500 Credits. Google Veo 3.1 Fast: high-speed text-to-video with native audio, up to 4K, supports 4/6/8-second clips.
- Veo 3.1 Image to Video API — video generation, 1,140 Credits. Google Veo 3.1 image-to-video: animate a still image into a cinematic clip with native audio, up to 4K, 4/6/8 seconds.
- Veo 3.1 Text to Video API — video generation, 1,140 Credits. Google Veo 3.1: flagship text-to-video with native audio, cinematic realism, up to 4K, 4/6/8-second clips.
- Wan 2.7 Image Pro API — image generation, 160 Credits. Tongyi Wanxiang Wan 2.7 Image Pro: text-to-image with up to 4K output, strong prompt adherence, and support for both Chinese and English prompts.
- Wan 2.7 Image-to-Video API — video generation, 334 Credits. Animate any still image into a high-quality video with natural motion, supporting first-frame, first+last-frame, and video continuation modes
- Wan 2.7 Text-to-Video API — video generation, 334 Credits. Alibaba's latest video generation model with cinematic quality, native audio support, and up to 1080P 15-second output
- wan3.0-video API — video generation, coming soon. Preview page for Alibaba Cloud Wan 3.0, an all-in-one video generation model currently in invite-only testing. HiAPI will add API access here after the production contract and availability are verified.
- Z-Image API — image generation, 16 Credits. Tongyi Z-Image: efficient photorealistic text-to-image, fast Turbo generation, accurate in-image text rendering in both Chinese and English.
Image to Image & Editing APIs
Black Forest Labs Model APIs