Frontier Models
xAI Grok Imagine Video 1.5 Adds Text-to-Video and Native 1080p
The model update introduces text-to-video generation, support for as many as seven reference images, and 1080p resolution options while rolling out to the API and consumer apps.
Grok Imagine Video 1.5 is xAI's AI video generation model updated to support text-to-video prompts, multiple reference images, and higher resolution output.
xAI has made Grok Imagine Video 1.5 generally available on the Imagine API under the identifier grok-imagine-video-1.5. The update has also been rolled out to users on grok.com/imagine as well as the iOS and Android applications. These changes build on the prior version by expanding input options and output quality. The model now handles additional generation modes while maintaining improvements in motion, physics, and audio.
What new modes of video generation does Grok Imagine Video 1.5 support?
The model now supports text-to-video generation in addition to image-to-video. This allows users to create videos directly from textual prompts without requiring an initial image input. The feature expands use cases for those who wish to visualize concepts from written descriptions alone. Reference-to-video functionality has been enhanced to accept up to seven reference images per request.
These references help ensure consistent characters, style, and visual elements throughout the generated video. Native 1080p resolution is supported for both text-to-video and image-to-video tasks on the grok-imagine-video-1.5 model. Reference-to-video remains capped at 720p resolution according to the developer documentation.
How does generation speed compare with earlier versions?
Grok Imagine Video 1.5 Fast produces 6-second 720p videos in about 25 seconds. This represents an improvement over the previous model's generation time of 40 or more seconds for similar outputs. The faster processing supports more rapid iteration during creative workflows. Users can access the Fast variant on the consumer platforms and through the API.
What audio and productivity features are included?
Videos produced by the model include automatically generated synchronized audio. This audio incorporates sound effects, ambience, and speech to match the visual content. The addition provides a more complete output without separate post-processing steps. New productivity features include Projects, multiple parallel agents, and library search.
These additions aim to enhance user workflow in the Grok platform. Developers can access the model through the xAI Imagine API for integration into applications. The rollout to consumer apps makes these features available to a wider audience. The combination of faster generation and higher resolution may lead to increased adoption in creative and professional settings.
How does this update affect the competitive landscape?
The enhancements position Grok Imagine Video 1.5 as a stronger option by providing more flexible input options and higher resolution outputs. The ability to use text prompts alone broadens accessibility for users without reference materials. Multiple reference images allow for better control over the output by providing examples of desired elements.
What do official announcements indicate about the changes?
The updates build on the previous launch by adding image and voice references along with prompt-based video and 1080p capabilities. Further details are available in the official documentation. The model supports a maximum of seven reference images per request for guiding the video generation process.
Today it goes further: image and voice references, video from a prompt alone, and native 1080p generation.xAI
What are the technical specifications for reference usage?
On the grok-imagine-video-1.5 model, users can provide a maximum of seven reference images per request. Text-to-video supports native 1080p while reference-to-video is limited to 720p. These specifications are detailed in the model capabilities documentation.
| Feature | Description | Limitation |
|---|---|---|
| Text-to-Video | Generates video from text prompt | Native 1080p supported |
| Image-to-Video | Generates video from image input | Native 1080p supported |
| Reference-to-Video | Uses up to 7 reference images | Capped at 720p |
| Audio | Synchronized sound effects and speech | Automatically generated |
| Speed (Fast) | 6-second 720p video | About 25 seconds |
- Review the model capabilities in the developer documentation.
- Prepare text prompts or reference images as needed.
- Submit the request via the API or app interface.
- Review the generated video including audio elements.
- Iterate using multiple reference images for consistency.
Frequently asked
What is the maximum number of reference images supported?
The Grok Imagine Video 1.5 model supports up to seven reference images per request for reference-to-video generation.
Which resolutions are available for different video types?
Text-to-video and image-to-video support native 1080p. Reference-to-video is limited to 720p.
Sources
- xAI — Grok Imagine Video 1.5 is now generally available on the Imagine API and rolled out to grok.com/imagine and apps, with better motion, physics, audio at fastest speeds.
- xAI — The update adds image and voice references, video from a prompt alone, and native 1080p generation.
- xAI — On grok-imagine-video-1.5, text-to-video supports native 1080p. Reference-to-video allows a maximum of 7 reference images per request.
- xAI — Imagine Video 1.5 brings text-to-video, image and voice references, plus native 1080p with more lifelike motion and sound.