MODELS · ONE SUBSCRIPTION, ALL OF THEM

Every model. One bill.

getvivix runs 234+ frontier AI models for video, image, and audio in a single studio. Pick any model to see what it does, or start free.

Start free

Video models · 85

Kling VIDEO 3.0 4K

True 4K Kling video in three fixed aspect ratios, with native audio and up to 6-shot sequencing

kling-ai
Kling VIDEO O3 4K

True 4K output in three aspect ratios with native audio and first/last-frame control, from $0.42/second

kling-ai
LTX-2 Fast

LTX-2's faster, lower-latency tier — 4K, 50fps, and optional synced audio

lightricks
LTX-2 Pro

Cinematic LTX-2 Pro text and image to video generator

lightricks
MiniMax H3

Omni-reference video generation with first/last frame control

minimax
PixVerse V6

Multi-shot cinematic video generation with native audio, 20+ camera controls, and character consistency

pixverse
Seedance 1.5 Pro

Native synced audio and video in one generation, with output up to 1080p

bytedance
Seedance 2.0

Premium multimodal video generation with native audio and cinematic motion

bytedance
Seedance 2.0 Fast

Seedance 2.0, tuned for speed — native audio and multi-shot motion intact, minus the flagship's 1080p tier

bytedance
Seedance 2.0 Mini

The lowest-priced tier of Seedance 2.0 — same 720p ceiling as Fast, native audio and multi-shot generation intact

bytedance
Veo 3.1 Fast

Google's speed-tuned Veo 3.1 model: generate from a prompt, an image, reference photos, or extend an existing clip

google
Wan2.7

The first Wan with reference-to-video and video-to-video editing — four modes, native audio, multi-shot scenes

alibaba
Aurora v1

24fps lip-sync avatar video from a single photo and an audio clip

creatify
Aurora v1 Fast

Aurora's faster variant: turns one photo and one audio clip into a 480p lip-synced avatar video.

creatify
FLUX 3 Video

Multimodal video generation with native synchronized audio across styles and modes

black-forest-labs
Gemini Omni Flash

Multimodal video generation and editing with native audio, references, and multi-turn control

google
Grok Imagine Video

AI video generation with synchronized audio from text and images

xai
Grok Imagine Video 1.5

Higher-tier Grok video from a start frame or reference images — now up to 1080p

xai
HappyHorse 1.0

Alibaba's video model with 4 modes — text, image, up to 9 references, or edit an uploaded clip, 720p/1080p, 3-15s

alibaba
HappyHorse 1.1

Alibaba's upgraded video model — text, image, or up to 9 reference images, 3-15s clips at 720p/1080p

alibaba
HeyGen Avatar IV

AI talking avatar video from a HeyGen avatar or your own photo, driven by a script or audio

heygen
HeyGen Avatar V

Script-to-video talking avatars from 500 ready HeyGen looks, no photo needed

heygen
HeyGen Video Agent

One prompt, a finished video: script, avatar, B-roll, and captions, up to 5 minutes long

heygen
Kling VIDEO 2.6 Pro

Kling VIDEO 2.6 Pro is a full audio-visual AI video model that combines cinematic-quality video generation with native audio (dialogue, sound effects, ambience), with optional Motion Control for precise character movement via the API.

kling-ai
Kling VIDEO 3.0 Pro

High-fidelity multimodal video generation with native audio and advanced editing

kling-ai
Kling VIDEO 3.0 Standard

Multimodal video generation with native audio and efficient performance

kling-ai
Kling VIDEO 3.0 Turbo

Fast Kling 3.0 video with native audio built in at no extra cost, up to 6 shots per generation

kling-ai
Kling VIDEO O3 Pro

Text-to-video, image-to-video, video editing, and motion transfer in one model — plus native audio.

kling-ai
Kling VIDEO O3 Standard

Kling 3.0's budget tier — 720p video with native audio and prompt-based edits

kling-ai
KlingAI 1.6 Pro

Turns a start and end frame into a 1080p video clip, 5 or 10 seconds long

kling-ai
KlingAI 1.6 Standard

Turn a text prompt or a single photo into a steady, accurate 720p video clip

kling-ai
KlingAI 2.0 Master

Anchor a first and last frame from your own photos, then render a 5 or 10 second 1080p clip

klingai
KlingAI 2.1 Master

1080p text- or image-to-video with 3-image character reference — the flagship Kling 2.1 tier

kling-ai
KlingAI 2.1 Pro

KlingAI 2.1 Pro for cinematic AI video generation

klingai
KlingAI 2.1 Standard

Animate a photo into 1080p video — the budget-friendly Kling 2.1 tier

kling-ai
KlingAI 2.5 Turbo Pro

Kling's higher tier: text and image to video in one model, up to 10-second clips, two guide frames

klingai
KlingAI 2.5 Turbo Standard

Kling 2.5 Turbo's budget tier — 720p image-to-video in 5 or 10 second clips

klingai
KlingAI Avatar 2.0 Pro

Photo plus audio becomes a 5-minute talking avatar video, no prompt required

kling-ai
KlingAI Avatar 2.0 Standard

Turns a photo and an audio track into a lip-synced avatar video

kling-ai
KlingAI Lip-Sync

Redub existing video to new dialogue or music, with independent original and synced audio volume control

kling-ai
LTX-2

Open-weight AI video model with synchronized audio, 20-second clips, and up to 120fps

lightricks
LTX-2 Retake

Redo one line, one shot, one sound cue — without touching the rest of the clip

lightricks
LTX-2.3

Multimodal video generation with native synchronized audio, text, image, or audio-driven input, and resolution up to 4K.

lightricks
LTX-2.3 Fast

Speed-optimized LTX-2.3 variant: text-to-video and image-to-video with native synced audio, up to 4K and 20-second clips.

lightricks
MiniMax 01 Director

Camera moves that hold steady, not random AI drift

minimax
MiniMax 01 Live

Animates static anime illustrations and manga panels into a steady 6-second, 720p character clip

minimax
MiniMax Hailuo 02

Two images become a video's first and last frame; 768p runs 10s, 1080p runs 6s

minimax
MiniMax Hailuo 2.3

High fidelity AI video generation from text or images

minimax
MiniMax Hailuo 2.3 Fast

MiniMax Hailuo 2.3 Fast animates one photo into a 6-10 second video clip

minimax
OmniHuman-1.5

AI talking avatar video from one photo and an audio clip

bytedance
P-Video

Fast, cheap AI video with synced dialogue

prunaai
P-Video-Animate

Animate one photo using the motion from an existing video.

prunaai
P-Video-Replace

Swap the character in any video — no reshoot needed

prunaai
PixVerse LipSync

Re-sync a video's mouth movements to your own audio track for dubbing and redubbing

pixverse
PixVerse V3.5

PixVerse V3.5 text-to-video with 20 built-in effect templates

pixverse
PixVerse V4

Animate photos or prompts with camera motion and 5 styles

pixverse
PixVerse V4.5

1080p text and image to video with camera movement control and start/end-frame guidance

pixverse
PixVerse V5

PixVerse V5 cinematic text to video and image to video

pixverse
PixVerse V5 Fast

Fast text to video and image to video generation for rapid iteration

pixverse
PixVerse V5.6

More accurate multi-character lip-sync, cleaner motion in busy scenes, optional native audio, up to 1080p video

pixverse
Runway Aleph 2.0

Localized video editing from a text prompt — swap products, backgrounds, or lighting without a reshoot, from $0.28/s

runway
Runway Gen-4 Turbo

Single-frame image-to-video, 5 or 10-second clips, no Runway subscription required

runway
Runway Gen-4.5

Text-to-video and image-to-video in five aspect ratios, with 5, 8, or 10 second clips

runway
Seedance 1.0 Pro

1080p text- and image-to-video with a 1.2 to 12-second range, plus a 480p mode to preview cheaply before the final render.

bytedance
Seedance 1.0 Pro Fast

Fast Seedance 1.0 Pro video generation for dance content

bytedance
Seedance 2.5

ByteDance's flagship video model — 30-second clips, 30 references, and prompt-based editing of an existing clip

bytedance
SkyReels V4

One video model that generates, edits, and extends footage, with 8-frame anchoring and synced audio up to 1080p

skywork
Sora 2

Next generation AI video and audio model from OpenAI

openai
Sora 2 Pro

OpenAI's 1080p Sora 2 tier — text-to-video or image-to-video, native audio, 4 to 20 second clips

openai
sync-3

Full-scene lip synchronization with global face understanding and obstruction handling

sync
Veo 2

Google's original Veo: camera-directed video with first-and-last-frame anchoring

google
Veo 3

Cinematic video generation, now with native audio

google
Veo 3 Fast

Google's faster, cheaper Veo 3 tier — with start-and-end frame anchoring flagship Veo 3 lacks

google
Veo 3.1

Veo 3.1 cinematic AI video with native audio

google
Vidu 2.0

1080p AI video with movement-amplitude control and built-in music

vidu
Vidu Q1

1080p AI video at 24fps, from a text prompt or first-and-last-frame images

vidu
Vidu Q2 Pro

AI video guided by 7 reference images

vidu
Vidu Q2 Turbo

Faster Q2-tier video with adjustable motion, built-in music, and 7 reference images

vidu
Vidu Q3

Multimodal video generation with native audio and intelligent shot planning

vidu
Vidu Q3 Turbo

Vidu's fast, low-latency video model with flexible sizing and optional audio

vidu
Wan2.2 A14B

MoE video generation from text or images at 480p to 720p

alibaba
Wan2.5-Preview

Alibaba's original Wan video model — 5-10s clips from text or an image, with native audio

alibaba
Wan2.6

Multimodal video generation with multi-shot and native sound

alibaba
Wan2.6 Flash

The fast, distilled Wan2.6 Flash — lip-synced or silent image-to-video, up to 15 seconds

alibaba
Wan3.0

Alibaba's higher-end multimodal video model — keyframes, references, and stronger character consistency

alibaba

Image models · 87

GPT Image 1.5

OpenAI's image model for text-to-image and editing — three sizes, four quality tiers, up to 16 reference images.

openai
GPT Image 2

OpenAI's newest image model — up to 16 reference images and native 4K output

openai
Grok Imagine Image Pro

Describe an edit in plain English — no masks, no sliders — then export up to 2K across nine aspect ratios

xai
Kling IMAGE 3.0

2K image generation from Kling AI, tuned for realistic textures and iterative image-to-image edits

kling-ai
Kling IMAGE O3

High-resolution text-to-image and editing with multi-reference consistency

kling-ai
Nano Banana 2

Google's Gemini 3.1 Flash Image — web-grounded search, 14 reference images, up to 2K output

google
Seedream 5.0 Lite

Reasoning-first Seedream 5.0 tier, pairing built-in reasoning with live web search for accurate, current images

bytedance
Seedream 5.0 Pro

ByteDance's flagship image model: precise masked edits, 1K or 2K output, and up to 10 reference images in one request

bytedance
Vivi

The getvivix signature model — instant images in any style

getvivix
Wan2.7 Image

Alibaba's unified image generator and editor — avatar customization, 9 reference images, and text in 12 languages

alibaba
Wan2.7 Image Pro

Alibaba pairs an optional thinking pass with 9-image reference grounding

alibaba
Bria 3.2

Text to image model trained on licensed stock from Getty, Alamy, and Envato

bria
Bria FIBO

Structured JSON prompts render reproducible images from licensed data

bria
Bria FIBO Edit

Edit photos from a text prompt — mask inpainting, outpainting, and generative fill

bria
Bria Fibo Edit Tools

Preset-driven photo editing — 7 structured operations from relighting to outpainting, no prompt writing needed

bria
DALL·E 2

DALL·E 2 online — no OpenAI account needed

openai
DALL·E 3

DALL·E 3 high fidelity text to image generation API

openai
Exactly Bold Chromatics

Vibrant, high-contrast illustrative style with bold color palettes

exactly
Exactly Bright Pulse

Bright, energetic photographic style with vivid lighting

exactly
Exactly Dark Comics

Heavy-shadow noir comic art, built for horror, thriller, and dystopian scenes

exactly
Exactly Distant Reality

Dreamy photographic style with surreal, distant atmosphere

exactly
Exactly Earthy Elegance

Warm, organic illustrative style with muted earth tones

exactly
Exactly Editorial Line

Clean, editorial-style line illustrations with refined detail

exactly
Exactly Extreme Contrast

High-contrast photographic style with dramatic light and shadow

exactly
Exactly Grain Film Look

Analog film photography style with natural grain and warm tones

exactly
Exactly Graphic Harmony

Balanced, harmonious graphic illustrations with cohesive composition

exactly
Exactly Graphic Novel

Comic book and graphic novel style with strong ink lines and dramatic shading

exactly
Exactly Graphite Creature

Graphite-shaded creature and character art, built for fur, scales, and muscle detail

exactly
Exactly Journey

Cinematic golden-hour travel tones for your reference photo, 11 ratios up to 2K

exactly
Exactly Monochrome Café

Monochromatic illustrative style with warm café-inspired tones

exactly
Exactly Muted Modern

Soft, desaturated illustration style built for modern branding and app visuals

exactly
Exactly Playful Line Adventures

Expressive, character-driven line art built for playful storybook and branding scenes

exactly
Exactly Warm Light

Soft, warm-lit photographic style with inviting golden tones

exactly
FLUX Virtual Try-On

Swap a garment onto any person photo in 4 fast steps, keeping the face, pose, logos, and stitching intact

black-forest-labs
FLUX.1 [dev]

12B open-weight model, second only to FLUX.1 [pro] in output quality

black-forest-labs
FLUX.1 [schnell]

Apache 2.0 open source FLUX.1 model with 1-4 step image generation

black-forest-labs
FLUX.1 Kontext [dev]

Open image editing model for fast iterative workflows

black-forest-labs
FLUX.1 Kontext [max]

Precision image editing, not text-to-image — up to 2 reference images per edit

black-forest-labs
FLUX.1 Kontext [pro]

Unifies text-to-image generation and precise editing in one model — 2 reference images, 9 aspect ratios up to 21:9

black-forest-labs
FLUX.1 Krea [dev]

Black Forest Labs' fine-tune of FLUX.1 dev, made with Krea AI to trade the oversaturated AI look for real photorealism

black-forest-labs
FLUX.1.1 [pro]

Black Forest Labs' well-rounded FLUX model for generating images and editing from a reference, in ten built-in aspect ratios.

black-forest-labs
FLUX.1.1 [pro] Ultra

Up to 4-megapixel FLUX images across nine aspect ratios, generated in about 10 seconds

black-forest-labs
FLUX.2 [dev]

The full 32B FLUX.2 model — precise control over prompts, references, and edits

black-forest-labs
FLUX.2 [flex]

The FLUX.2 tier that exposes manual sampling controls and up to 10 reference images for precise edits

black-forest-labs
FLUX.2 [klein] 4B

Apache 2.0 open license — the fastest Klein model for real-time editing

black-forest-labs
FLUX.2 [klein] 4B Base

Full 28-step image generation and editing from Black Forest Labs' undistilled klein foundation checkpoint

black-forest-labs
FLUX.2 [klein] 9B Base

Undistilled foundation model for high-quality image generation and editing

black-forest-labs
FLUX.2 [klein] 9B KV

KV-cache accelerated image generation and editing for real-time multi-reference workflows

black-forest-labs
FLUX.2 [max]

Black Forest Labs' FLUX.2 tier that grounds images in live web search, with 8-image reference editing

black-forest-labs
FLUX.2 [pro]

High control FLUX.2 Pro image generation and editing

black-forest-labs
GPT Image 1

GPT Image 1 high fidelity image generation for GPT-4o

openai
Grok Imagine Image

xAI's Grok Imagine model for text-to-image and image-to-image, the standard tier in its family

xai
Grok Imagine Image Quality

xAI's quality-focused image generation and editing — sharper realism, better text rendering, tighter prompt following

xai
HiDream-I1 Dev

HiDream's balanced 28-step tier — quicker than the 50-step MIT-licensed Full, more steps than Fast

runware
HiDream-I1 Fast

HiDream-I1 Fast for low latency text to image generation

runware
HiDream-I1 Full

MIT-licensed, 50-step flagship — the un-distilled source that Dev and Fast are distilled from

runware
Ideogram 2.0

Sharp typography from Ideogram's per-style rendering model, released August 2024

ideogram
Ideogram 3.0

Ideogram 3.0 text to image model for sharp design visuals

ideogram
Ideogram 4.0

Ideogram's first open-weight model — precise typography, structured layout control, and transparent 2K output

ideogram
Imagen 3

Google's photorealistic Imagen 3 model for detailed, high fidelity images

google
Imagen 3 Fast

Google's speed-tuned Imagen 3 variant, built for low-latency image generation

google
Imagen 4 Fast

High speed Imagen 4 Fast text to image generation

google
Imagen 4 Preview

High fidelity 2K text to image generation by Google

google
Imagen 4 Ultra

Google's most prompt-accurate Imagen model, for photorealistic detail and sharp text

google
ImagineArt 1.5 Pro

Professional AI image generation with native 4K and refined visual control

imagineart
ImagineArt 2.0

Reasoning-based text to image generation with vibrant true-to-life color

imagineart
Juggernaut Lightning Flux by RunDiffusion

A fast, low-cost Flux checkpoint for quick drafts and high-volume runs.

rundiffusion
Juggernaut Pro Flux by RunDiffusion

Photorealistic Flux checkpoint, hosted and ready — no download, no GPU, no setup.

rundiffusion
Kandinsky 5.0 Image Lite

Open-source text-to-image generation and image editing, built for fast, prompt-accurate 1K output.

runware
Krea 2 Large

Krea 2's biggest, rawest tier — over 2x Medium's size for stronger photorealism, film grain, and texture

krea
Krea 2 Medium

Krea's smaller, faster tier — under half the size of Large, tuned for stable illustration

krea
Nano Banana

Gemini 2.5 Flash Image — the original Nano Banana, built for consistent-likeness photo edits

google
Nano Banana 2 Lite

Lighter Nano Banana 2 image model for faster generation and editing workflows

google
P-Image

Pruna AI's real-time text-to-image model — sub-second generation across nine size presets, square to 21:9 ultrawide

prunaai
P-Image-Edit

Edit with up to 4 reference images — composition and style stay locked

prunaai
Qwen-Image

Apache 2.0 open source foundation model built for precise Chinese and English text in images

alibaba
Qwen-Image-2.0

Alibaba's lighter next-gen model for clean in-image text and layouts

alibaba
Qwen‑Image‑Edit

Edit an image with a text instruction, including bilingual Chinese-English text swaps that keep the font

alibaba
Recraft V4

Professional text-to-image model for brand and marketing design

recraft
Recraft V4 Pro

Advanced design-focused image generation with enhanced control and fidelity

recraft
Reve 2.1

Native 4K x 4K, 16-megapixel image generation with layout-aware, region-by-region editing

reve
Seedream 4.0

One model for text-to-image generation and natural-language editing, up to 4K with 14-image reference fusion

bytedance
Stable Diffusion 3

From Stability AI, live here now — legible in-image text, multi-subject scenes, free to try

stabilityai
Wan2.5-Preview Image

Wan2.5's still-image model — one generated frame, not a video clip

alibaba
Wan2.6 Image

Edit photos with up to four reference images, or generate from a prompt, on Alibaba's Wan2.6 stack

alibaba
Z-Image

Apache 2.0 open-source model from Alibaba — the full 6B Z-Image, built for creative control, not just speed

alibaba
Z-Image-Turbo

Fast photorealistic image generator with text control

alibaba

Audio models · 19

ACE-Step v1.5 Base

Open-source (MIT) music generation with editable lyrics, song covers, and 50+ languages

runware
ACE-Step v1.5 Turbo

Turns a lyric sheet into a full track faster — fewer denoising steps, same engine as the Base model.

runware
Eleven Flash v2

English-only TTS at about 75ms latency, built for live streams, games, and voice apps

elevenlabs
Eleven Flash v2.5

ElevenLabs' fast, affordable TTS model — 75ms latency, 32 languages, built for voice agents and bulk narration

elevenlabs
Eleven Monolingual v1

The original Eleven TTS, English-only

elevenlabs
Eleven Multilingual v1

Legacy multilingual TTS across 9 languages

elevenlabs
Eleven Multilingual v2

ElevenLabs' expressive, 29-language voice model — built for emotional nuance over raw speed, no download required

elevenlabs
Eleven Music v1

Describe a song and get a finished track back — vocals or instrumental, structured section by section

elevenlabs
Eleven Turbo v2

First-generation low-latency Eleven TTS model, English only

elevenlabs
Eleven Turbo v2.5

First-generation low-latency Eleven TTS model, 32 languages

elevenlabs
Eleven v3

Dramatic multi-speaker dialogue TTS, 70+ languages

elevenlabs
Gemini 3.1 Flash TTS

Expressive text-to-speech with audio tags, multi-speaker dialogue, and 70+ languages

google
Inworld TTS-1.5 Max

Expressive, broadcast-ready text-to-speech: 73 voices, 15 languages, sub-250ms response time

inworld
Inworld TTS-1.5 Mini

Sub-130ms text-to-speech powering getvivix's avatars, agent, and chat voices

inworld
MiniMax Speech 2.8

332 voices across 24 languages, with live emotion control and up to 50,000 characters per script

minimax
Qwen3-TTS 1.7B Base

Clones a voice from a 3-second sample, then generates natural multilingual speech in 10+ languages at 97ms latency

alibaba
Qwen3-TTS 1.7B CustomVoice

Nine named preset voices with natural-language style control

alibaba
Qwen3-TTS 1.7B VoiceDesign

Text-to-speech that designs a brand-new voice from your written description, no sample audio required

alibaba
xAI Text-to-Speech

Text-to-speech with six voices, inline speech tags for pauses and emphasis, and 20 supported languages

xai

Text models · 24

Claude Fable 5

Anthropic's top Claude model — Fable 5 sits above Opus in capability, with 1M-token context, vision, and adaptive thinking

anthropic
Claude Haiku 4.5

Anthropic's fastest, most cost-efficient Claude tier — 200K context, 64K output tokens, built for sub-agent dispatch

anthropic
Claude Opus 4.7

Free to start, 100+ other AI models on one account — Anthropic's top Claude 4.7, 1M-token context, xhigh thinking tier

anthropic
Claude Sonnet 4.6

Anthropic's daily-driver Claude — reads images, six adjustable thinking levels, and 65,536-token replies

anthropic
DeepSeek V4 Flash

Open-weight MoE model, MIT-licensed, with a 1M-token context and switchable reasoning depth

deepseek
Gemini 3 Flash

Multimodal understanding model for text, image, video, and audio with a 1M-token context

google
Gemini 3.1 Flash Lite

Google’s fast, low-cost Gemini 3.1 tier for high-volume translation, transcription, and document extraction

google
Gemini 3.1 Pro

Google's multimodal reasoning model with a 1M-token context window

google
GLM-4.7

Z.ai's open-weight value LLM — 200K context, 73.8% on SWE-bench Verified

zai
GLM-5.1

Z.ai's open-weight flagship LLM — 200K context, MIT license, 8-hour autonomous agent runs

zai
GPT-5.4

OpenAI's flagship model — 1M-token context, native computer-use, and 33% fewer factual errors than GPT-5.2

openai
GPT-5.4 Mini

OpenAI's compact GPT-5.4 sibling: 400K context, vision input, and computer use for coding agents and subagent workflows

openai
GPT-5.4 Nano

Ultra-low-latency LLM for high-volume classification, extraction, and lightweight automation

openai
GPT-5.5

OpenAI's newest flagship LLM — deepest reasoning, computer-use, 1M+ context

openai
Kimi K2.6

262K context, configurable reasoning, and an 80.2 SWE-Bench Verified score

moonshotai
LLaVA-1.6-Mistral-7B

Turns an image into a caption, guided by an optional prompt from a Mistral 7B backbone

runware
MiniMax M2.5

MiniMax's original agentic-coding and office-task model — released before the longer-context M2.7 and faster M2.7 Highspeed

minimax
MiniMax M2.7

MiniMax's 230B-parameter agent model also drives an AI Character's dialogue on getvivix

minimax
MiniMax M2.7 Highspeed

Faster throughput for agentic coding and tool‑driven automation

minimax
Open Age Detection

Guess your age with AI

runware
OpenAI CLIP ViT-L/14

A vision-language model that reads an image and writes a caption for it

openai
Qwen2.5-VL-3B-Instruct

Caption or answer questions about any photo instantly — no download, GPU, or license required

alibaba
Qwen2.5-VL-7B-Instruct

Open-weight vision-language model that reads and describes images, charts, and documents

alibaba
ViT Age Classifier

Vision Transformer model that classifies facial age into brackets with a confidence score

runware

Utility models · 14

3D models · 5