Saturday, August 1, 2026

Today’s Edition

AI Intel Report

MARKETS

Frontier Models

xAI Grok Imagine Video 1.5 Adds Text-to-Video and Native 1080p

The model update introduces text-to-video generation, support for as many as seven reference images, and 1080p resolution options while rolling out to the API and consumer apps.

4 MIN READ
Inside a contemporary xAI research laboratory a wide angle view captures a long minimalist workstation table made of brushed aluminum holding three large flat panel displays arranged side by side each monitor displaying crisp high definition video frames without any visible lettering or icons the central screen shows a smooth 1080p video sequence of a city street at dusk with moving vehicles and pedestrians captured at native resolution the left screen presents seven separate reference photographs arranged in a clean grid depicting various angles of the same urban environment including close ups of architecture vehicles and lighting conditions the right screen displays an additional set of reference images focused on human figures and motion paths all images rendered in photorealistic detail the workstation surface also contains a wireless keyboard a precision mouse several external hard drives connected via visible cables and a notebook with blank pages a person wearing a plain dark sweater sits with their back to the camera facing the monitors their hands resting on the keyboard in a posture suggesting active interaction with the system behind the workstation a glass wall reveals a server room filled with tall black server racks and blinking indicator lights on networking equipment the floor is covered with neutral gray carpet and the walls feature subtle acoustic panels the overall environment is brightly lit by overhead fluorescent panels creating even illumination across the workspace no logos badges or text appear anywhere in the scene the composition emphasizes the hardware setup supporting advanced video generation capabilities including multi image reference handling and full 1080p output the scene remains entirely static and documentary in nature focusing on the physical arrangement of computing equipment reference materials and the laboratory interior associated with xAI development of Grok Imagine Video 1.5 the camera position is slightly elevated to include both the workstation and the distant server racks providing context for the computational infrastructure required for text to video workflows and API rollout the entire view stays grounded in real world office and data center elements without any stylized additions or overlays
Illustration: AI Intel Report

Grok Imagine Video 1.5 is xAI's AI video generation model updated to support text-to-video prompts, multiple reference images, and higher resolution output.

xAI has made Grok Imagine Video 1.5 generally available on the Imagine API under the identifier grok-imagine-video-1.5. The update has also been rolled out to users on grok.com/imagine as well as the iOS and Android applications. These changes build on the prior version by expanding input options and output quality. The model now handles additional generation modes while maintaining improvements in motion, physics, and audio.

What new modes of video generation does Grok Imagine Video 1.5 support?

The model now supports text-to-video generation in addition to image-to-video. This allows users to create videos directly from textual prompts without requiring an initial image input. The feature expands use cases for those who wish to visualize concepts from written descriptions alone. Reference-to-video functionality has been enhanced to accept up to seven reference images per request.

These references help ensure consistent characters, style, and visual elements throughout the generated video. Native 1080p resolution is supported for both text-to-video and image-to-video tasks on the grok-imagine-video-1.5 model. Reference-to-video remains capped at 720p resolution according to the developer documentation.

How does generation speed compare with earlier versions?

Grok Imagine Video 1.5 Fast produces 6-second 720p videos in about 25 seconds. This represents an improvement over the previous model's generation time of 40 or more seconds for similar outputs. The faster processing supports more rapid iteration during creative workflows. Users can access the Fast variant on the consumer platforms and through the API.

What audio and productivity features are included?

Videos produced by the model include automatically generated synchronized audio. This audio incorporates sound effects, ambience, and speech to match the visual content. The addition provides a more complete output without separate post-processing steps. New productivity features include Projects, multiple parallel agents, and library search.

These additions aim to enhance user workflow in the Grok platform. Developers can access the model through the xAI Imagine API for integration into applications. The rollout to consumer apps makes these features available to a wider audience. The combination of faster generation and higher resolution may lead to increased adoption in creative and professional settings.

How does this update affect the competitive landscape?

The enhancements position Grok Imagine Video 1.5 as a stronger option by providing more flexible input options and higher resolution outputs. The ability to use text prompts alone broadens accessibility for users without reference materials. Multiple reference images allow for better control over the output by providing examples of desired elements.

What do official announcements indicate about the changes?

The updates build on the previous launch by adding image and voice references along with prompt-based video and 1080p capabilities. Further details are available in the official documentation. The model supports a maximum of seven reference images per request for guiding the video generation process.

Today it goes further: image and voice references, video from a prompt alone, and native 1080p generation.xAI

What are the technical specifications for reference usage?

On the grok-imagine-video-1.5 model, users can provide a maximum of seven reference images per request. Text-to-video supports native 1080p while reference-to-video is limited to 720p. These specifications are detailed in the model capabilities documentation.

Key capabilities and limits of Grok Imagine Video 1.5
FeatureDescriptionLimitation
Text-to-VideoGenerates video from text promptNative 1080p supported
Image-to-VideoGenerates video from image inputNative 1080p supported
Reference-to-VideoUses up to 7 reference imagesCapped at 720p
AudioSynchronized sound effects and speechAutomatically generated
Speed (Fast)6-second 720p videoAbout 25 seconds
  1. Review the model capabilities in the developer documentation.
  2. Prepare text prompts or reference images as needed.
  3. Submit the request via the API or app interface.
  4. Review the generated video including audio elements.
  5. Iterate using multiple reference images for consistency.

Frequently asked

What is the maximum number of reference images supported?

The Grok Imagine Video 1.5 model supports up to seven reference images per request for reference-to-video generation.

Which resolutions are available for different video types?

Text-to-video and image-to-video support native 1080p. Reference-to-video is limited to 720p.

Sources

  1. xAI — Grok Imagine Video 1.5 is now generally available on the Imagine API and rolled out to grok.com/imagine and apps, with better motion, physics, audio at fastest speeds.
  2. xAI — The update adds image and voice references, video from a prompt alone, and native 1080p generation.
  3. xAI — On grok-imagine-video-1.5, text-to-video supports native 1080p. Reference-to-video allows a maximum of 7 reference images per request.
  4. xAI — Imagine Video 1.5 brings text-to-video, image and voice references, plus native 1080p with more lifelike motion and sound.