Generate videos from text prompts, animate still images, or edit existing videos with natural language. The API supports configurable duration, aspect ratio, and resolution for generated videos — with the SDK handling the asynchronous polling automatically.
Try Grok Imagine Video with a real request, review its parameters and validate the output before integrating the API.
Use the quick estimate above for a single run, then review the full CometAPI and official-price comparison.
Explore competitive pricing for Grok Imagine Video, designed to fit various budgets and usage needs. Our flexible plans ensure you only pay for what you use, making it easy to scale as your requirements grow. Discover how Grok Imagine Video can enhance your projects while keeping costs manageable.
| Category | Item | Price |
|---|---|---|
| Input Pricing | Text | N/A (Free) |
| Image | $0.0016 | |
| Video per second | $0.008 | |
| Output Pricing | 480p | $0.04 |
| (Per second by resolution) | 720p | $0.056 |
Note: When generating video via API, you are charged per second. You will also be charged when using video or images as input.
Copy a working endpoint and code example, then open the complete API reference when you need every parameter.
Authenticate once, call the model endpoint and keep the same billing and observability workflow across providers.
Access comprehensive sample code and API resources for Grok Imagine Video to streamline your integration process. Our detailed documentation provides step-by-step guidance, helping you leverage the full potential of Grok Imagine Video in your projects.
Scan the model facts that matter before you choose an architecture or estimate production workload.
Use Grok Imagine Video for production workflows that match its video capabilities, then compare alternatives before committing to a long-term integration.
Text-to-video concepts and social content
Product demos and cinematic storyboards
Image-to-video motion and creative iteration
Compare other models available through CometAPI for different quality, latency, capability and pricing trade-offs.
Super powerful video generation model, with sound effects, supports chat format.
Sora 2 Pro is our most advanced and powerful media generation model, capable of generating videos with synchronized Audio. It can create detailed, dynamic video clips from natural language or images.
Midjourney video generation
What is Veo 3.1-Fast Veo 3.1-Fast is Google’s speed-optimized variant of the Veo 3.1 family of generative video models. It is explicitly tuned to reduce latency and cost for short, social-length video generation while preserving the improved audiovisual fidelity introduced in Veo 3.1. Veo 3.1 and Veo 3.1-Fast add richer native audio generation, stronger prompt adherence, and new editing flows (for example: first/last-frame interpolation, “ingredients to video”, and scene extension) compared with previous Veo releases.
Identify Face
Creating a Digital Human Task
Review live heartbeat data, endpoint availability and observed response times before moving into production.