AI Video APIs
AI Video API Collections
Aggregate Sora, Veo, and more with one API. Enjoy up to 80% discounts, transparent billing (no charges for blocked prompts), and stable, cinema-grade generation.
Available Video APIs
Veo 3.1 Official
Google Veo 3.1 Official on Vidgo API with Lite, Fast, and Quality tiers for text-to-video, image-to-video, first/last-frame, reference, native audio, and 4K workflows.
per second | Lite starts at 3.6 credits/s ($0.018), Fast starts at 10 credits/s ($0.050), Quality starts at 24 credits/s ($0.120)
Happy Horse
Alibaba Happy Horse video generation and editing on Vidgo API with text-to-video, image-to-video, reference-to-video, and video-edit workflows.
per second | generation/reference: 720p 16 credits/s, 1080p 32 credits/s | video edit: 720p 32 credits/s, 1080p 64 credits/s | about 43% lower than Fal.ai
Seedance 2
ByteDance's multimodal video generation family on Vidgo API with prompt-led creation, first/last frame guidance, reference image/video/audio control, optional native audio, and pay-per-second pricing.
per second with video input on Seedance 2 Fast, up to 40 credits/s on Seedance 2
Seedance 2.0
ByteDance Seedance 2.0 Standard and Fast video generation with text, ordered start/end frames, or multimodal image, video, and audio references.
per output second; reference-video workflows use separate discounted rates and billed reference-video seconds
Seedance 2.5
ByteDance Seed's Seedance 2.5 video model with separate text-to-video, start/end-frame image-to-video, and multimodal reference-to-video API workflows at 480p or 720p for 4 to 30 seconds.
per output second at 480p; other resolutions and reference-video billing use separate rates
Seedance 2.0 Mini
ByteDance's Seedance 2.0 Mini provides text-to-video, ordered start/end-frame image-to-video, and image, video, or audio reference-to-video modes at 480p or 720p.
per output second | 480p: 10 credits/s, 720p: 24 credits/s; reference-video workflows: 480p 6 credits/s, 720p 12.5 credits/s for output plus separately floored reference-video seconds
MiniMax H3
MiniMax H3 on Vidgo API with separate text-to-video, start/end-frame image-to-video, and multimodal reference-to-video workflows at fixed 2K output.
per generated second | reference video seconds are also billed at 21 credits/s; reference images after the first five add 6.4 credits each
Sora 2 Official
Sora 2 Official with synced audio, improved physics, optional reference image input, fixed 4-second to 20-second tiers, and Pro Official resolution-based output.
per 4s | 8s: 96 credits | 12s: 144 credits | 16s: 192 credits | 20s: 240 credits | $0.06 per second
Kling 3.0 Motion Control
Kling 3.0 Motion Control is a reference-driven motion transfer model that combines one character image and one source video with transparent per-second pricing.
per second at 720p | 1080p: 15 credits per second | 65% cheaper than fal.ai.
Kling 2.6 Motion Control
Kuaishou's motion control model that transfers motion from reference videos to character images while maintaining identity and adapting environments.
per second (720p), 12 credits per second (1080p)

Wan 2.6
Alibaba's Wan 2.6 video generation family for text-to-video, image-to-video, and video-to-video with multi-shot 1080p output.
per 5s (720p), 10s/160, 15s/240, 1080p: 5s/120, 10s/240, 15s/360, 20% cheaper than fal
Hailuo 02
MiniMax's #2 globally-ranked video model with NCR architecture, ultra-realistic physics, and 1080p cinematic output.
per second (768p), hailuo-02-pro 65 credits fixed
Seedance 1.0 Pro
ByteDance's #1 ranked video model with multi-shot storytelling, cinema-grade motion, and bilingual text-to-video generation.
per 5s (720p), 43cr (1080p), 10s doubles
Kling 2.6
Kuaishou's revolutionary video model that simultaneously generates visuals with synchronized dialogue, sound effects, and ambient audio in one pass.
and 10s/130 credits, 5s+audio/120 credits, 10s+audio/240 credits . 15% cheaper than Fal.ai

Wan Animate
Alibaba's 14B-parameter character animation model that transfers motion from reference videos to static characters with exceptional identity preservation.
per second, 480p/7 credits per second, 580p/12 credits per second, 720p/15 credits per second
Sora 2
OpenAI Sora 2 with Cameo, synchronized dialogue, and C2PA provenance.
per generation | sora-2: 0.15$ | sora-2-pro: 0.5$ | about 85% cheaper than replicate or fal.ai
Sora 2 Pro
Sora 2 Pro for 1024p fidelity, richer motion, and advanced physics accuracy.
per generation
Veo 3.1
Google Veo 3.1 for 1080p, 60s video with ingredients-to-video and frames-to-video.
veo3.1-lite: 720p/1080p 20cr, 4K 30cr; veo3.1-fast: 720p/1080p 36cr, 4K 60cr; veo3.1-quality: 720p/1080p 200cr, 4K 400cr
Hailuo 2.3
MiniMax's Hailuo 2.3 video model for realistic human motion, expressive characters, and text-to-video or first-frame guided generation at 768p and 1080p.
per 6s at 768p | 10s at 768p: 70 credits | 6s at 1080p: 60 credits | 37.5% cheaper than fal.ai.
Kling 1.6
Cost-effective Kling 1.6 access on Vidgo API for realistic video generation, text-to-video, image-to-video, Pro first/last-frame control, and Elements reference-image workflows.
per second for Standard | Pro: 15 credits/sec | 20% cheaper than fal.ai.
Kling 2.1
Kling 2.1 on PoYo provides Standard and Pro image-to-video modes with 5-second and 10-second clips, start-frame control, and optional end-frame control in Pro.
per 5s standard | pro starts at 55 credits | 50% cheaper than fal.ai.
Kling 2.5 Turbo Pro
Kling 2.5 Turbo Pro is a flexible short-form video model with text-to-video, optional frame guidance, smooth motion, cinematic depth, and fixed 5-second and 10-second tiers.
per 5s | 10s: 84 credits | 40% cheaper than fal.ai.

Wan 2.2 Fast
Wan 2.2 Fast provides fast text-to-video and image-to-video generation with low-cost 480p and 720p tiers for quick iteration.
per generation at 480p | 720p: 12 credits | 40% cheaper than official

Wan 2.5
Wan 2.5 combines text-to-video and image-to-video generation with 5-second and 10-second output, synchronized audio support, and multiple size and resolution tiers.
starts at 30 credits for 5s | 10s tiers available | 40% cheaper than fal.ai
Runway Gen-4.5
Runway Gen-4.5 is a high-fidelity video model focused on prompt adherence, cinematic motion, visual fidelity, and optional reference image guidance.
per 5s | 10s: 150 credits | 30% cheaper than official
Kling 3.0
Kuaishou's most advanced video model with native 4K/60fps output, multi-shot storyboarding, multilingual audio, and character consistency for up to 3 people.
| kling-3.0/standard no audio: 27 credit/s , with audio: 39 credit/s | kling-3.0/pro no audio: 39 credit/s , with audio: 49 credit/s | 10-20 % cheaper than Fal.ai
Seedance 1.5 Pro
ByteDance's latest video model with synchronized audio generation, flexible aspect ratios, and enhanced motion control.
480p: 4s/9cr, 8s/18cr, 12s/21cr; 720p: 4s/16cr, 8s/32cr, 12s/42cr; with audio doubles; 24-30% cheaper than replicate or fal.ai
Grok Imagine Video 1.5
Generate 1–15 second videos from text, one first-frame image, or up to seven reference images, with output up to 1080p.
per second
Grok Imagine
xAI's Aurora-powered visual AI for image generation and video creation with Fun, Normal, and Spicy creative modes.
per image, grok-imagine 30 credits per video
Comparison
vidgo.ai vs FAL vs Replicate vs KIE
Compare leading AI video generation platforms for production teams.
| Feature | vidgo.ai | FAL | Replicate | KIE |
|---|---|---|---|---|
| Discounts | Up to 80% OFF | Standard | Standard | Standard |
| Pricing Transparency | Credits; blocked prompts not charged | Per-call | Per-call | Per-call |
| Stability | Production-grade, steady QoS | Varies by model | Varies by provider | Varies by provider |
| API Access | Full REST + docs | REST | REST | REST |
| Playground Support | Support | Support | Support | Support |
FAQ
Frequently Asked Questions
Answers to common questions about the AI Video API.