Veo 3 is Google DeepMind's video model: you type a scene, or hand it a starting image, and it returns a short clip with the sound already in it. Google released Veo 3 on May 20, 2025 and Veo 3.1 on October 15, 2025, and Veo 3.1 is the version you can actually use in September 2026. Here is what it makes, the limits that shape every clip, how it ranks against other models, and why the Gemini app now makes videos with a different model.
What is Veo 3.1?
Veo 3.1 is the current release of Google's Veo model family. You describe a scene, or give it a first frame, and it generates a clip of 4, 6 or 8 seconds in which the picture, the dialogue, the sound effects and the background noise are all made in the same pass. Veo 3 introduced that joint audio. Veo 3.1 improved it and added more control over what the clip shows.
| Version | Released | What it added |
|---|---|---|
| Veo 3 | May 20, 2025 | The first Veo model to generate audio with the video, including dialogue with lip sync. It launched in the Gemini app and Flow for Ultra subscribers in the US. |
| Veo 3.1 | October 15, 2025 | Richer audio, stronger prompt adherence when turning an image into video, and sound in Flow's Ingredients, Frames and Extend features. |
| Veo 3.1 Lite | Preview in the Gemini API (updated March 2026) | A lower-cost version that tops out at 1080p. In the API it takes no reference images and can't extend clips. |
The original Veo 3 models are now marked deprecated in Google's API documentation. So when you search "what is Veo 3" today, the model you can open and use is Veo 3.1.
What Veo 3.1 can do
Veo 3.1 makes one short clip per request, from text alone or from text plus images. As of September 2026, Google's Gemini API documentation lists these specs:
| Spec | Veo 3.1 |
|---|---|
| Inputs | A text prompt, plus an optional first frame, an optional last frame (only with a first frame), and up to 3 reference images |
| Clip length | 4, 6 or 8 seconds; must be 8 for 1080p, 4K, reference images or extension |
| Resolution | 720p by default; 1080p and 4K |
| Shape | Wide 16:9 (the default) or vertical 9:16 |
| Frame rate | 24 fps |
| Sound | Always on: dialogue, sound effects and ambient sound |
| Longer videos | Extend by 7 seconds at a time, up to 20 times, for up to 148 seconds |
| Watermark | SynthID, an invisible AI watermark |
Three of those inputs change what you can control:
- First and last frame. Give Veo a starting image and an ending image, and it generates the motion between them. A product photo can open the clip and a close-up of the label can end it.
- Reference images. Up to 3 images keep a character, an object or a style the same across clips. Flow calls these "Ingredients" and recommends product references on a plain background.
- Extension. Each extension continues from the last second of the previous clip, so a series of 8-second shots can become one longer scene. In Flow, the extension always runs on Veo 3.1 Lite.
Sound follows your wording. Google's prompt guide says to put spoken lines in quotation marks, name each sound effect outright ("tires screeching loudly, engine roaring") and describe the soundscape of the place. Written that way, one line of a cafe scene reads: A barista slides a cup across the counter and says, "Oat latte for Sam," while the espresso machine hisses behind her.
Where Veo 3.1 falls short
Veo 3.1's limits are mostly about length, language and people:
- Eight seconds per generation. Anything longer is built from extensions or from separate clips joined in an editor.
- English first. Google fully supports English prompts and hasn't evaluated other languages, so dialogue in Spanish or Japanese may come out wrong.
- Short speech. Google's Veo page lists natural, consistent delivery of brief spoken lines as unfinished work. Check lip sync on quick lines.
- People rules. In the API, clips that start from an image or use reference images can show adults only, and in the EU, the UK, Switzerland and MENA every request is limited to adults.
- Blocked clips. Veo can refuse a generation for safety or audio-processing reasons. Google doesn't charge for a blocked video.
- Files expire. Clips made through the API are removed from Google's servers after 2 days, so download them promptly. A request takes from 11 seconds to 6 minutes at peak times.
On blind votes, newer models have pulled ahead of Veo 3.1. Artificial Analysis collects votes between pairs of unlabeled clips, sound included, and converts them into an Elo rating for each model. As of September 28, 2026, Veo 3.1 scores 1088 on text to video (14th) and 1082 on image to video (11th), while Google's newer Gemini Omni Flash leads text to video at 1233 and ByteDance's Seedance 2.0 scores 1210.
What is Veo 3.1 in Gemini?
In the Gemini app, video now comes from Gemini Omni, not Veo. Google's Gemini Omni page says Omni replaces Veo in the Gemini app, and the Gemini Apps help page names Omni as its video model. An Omni video is 10 seconds long with sound, made from a prompt, up to 5 photos or a single uploaded video, and you refine it by replying in the chat. On a personal account it requires a Google AI plan (the lowest is Google AI Plus at $4.99 a month), and users under 18 can't use it.
Veo 3.1 hasn't disappeared from Google's plans. As of September 2026, the Google AI plans page lists a limited Veo 3.1 Lite trial on Google AI Pro and "video generation with Veo 3.1" on Google AI Ultra. To choose the Veo model yourself, use Google Flow. The steps for using Veo 3.1 without an Ultra plan walk through Flow, the Gemini app and the API.
Where you can use Veo 3.1
Google Flow is the main place to use Veo 3.1 without writing code. As of September 2026:
| Where | Who it suits | Cost to start |
|---|---|---|
| Google Flow, on the web and in its mobile app | Anyone 18 or older in over 150 countries and territories | 50 free credits a day; a Veo 3.1 Lite clip costs 10 credits and a Veo 3.1 Quality clip costs 100 |
| Gemini API and Google AI Studio | Developers building clips into an app | Billed per second of video, with no free tier |
| Toybox AI's Video Creator | People who make images and flyers in the same account | Currently 350 credits per clip, on a Pro plan that costs $9.99 a month for 1,000 credits |
Google also lists Google Vids among the places to try Veo. Flow's plan credits come on top of the daily 50: Google AI Plus adds 200 a month for $4.99, Pro adds 1,000 for $19.99, and Ultra adds 10,000 or 25,000 on its $99.99 and $199.99 plans. Unused credits don't roll over, and Google may pause video for accounts without a plan during peak hours, around 2 PM to 5 PM UTC. The Veo 3.1 pricing breakdown turns those credits into a price per clip.
How Veo 3.1 compares with Sora 2 and Kling 3.0
OpenAI's Sora 2 is no longer an option. OpenAI discontinued the Sora app on April 26, 2026 and the Sora API on September 24, 2026. On the same Artificial Analysis board, the December version of Sora 2 scored 1083 against Veo 3.1's 1088, which is within the margin of error. The Sora vs Veo comparison covers how the two differed while both were live.
Kling 3.0 is the closer comparison now. It makes 3 to 15 seconds in one generation, can cut between shots inside a clip, and voices dialogue in Chinese, English, Japanese, Korean and Spanish. On votes it is level with Veo 3.1 on text to video, at 1095 for its 1080p version. The Veo 3 vs Kling breakdown compares plans and prices, and the roundup of the best AI video generators ranks the wider field.
How to describe a shot for Veo 3.1
Google's prompt guide builds a Veo prompt from three required parts and a few optional ones:
- Subject. Who or what is on screen: "a gray tabby cat", "a red bicycle leaning on a brick wall".
- Action. What happens in the clip: "stretches and yawns on a sunny windowsill".
- Style. The look: "soft natural light, shallow depth of field".
- Camera and composition (optional). "Slow push-in at eye level", "wide shot from across the room".
- Ambiance (optional). The light and mood of the place: "early morning, quiet kitchen".
Plan one shot per clip. For 8 seconds, write one subject, one action and one camera move, and split a story with three beats into three clips or a clip plus two extensions.
Make a Veo 3.1 clip in Video Creator
Video Creator in Toybox is powered by Google Veo 3.1 and turns a description, with an optional start photo, into one clip.
- Write the shot in "Describe your video" (up to 2,000 characters), in Google's subject, action, style order.
- To start from your own image, tap "Start from a photo". Adding an "End frame" fixes the last shot and locks the length at 8 seconds.
- Under "Format", pick Wide (16:9) for YouTube or Vertical (9:16) for Reels and TikTok, then pick 4, 6 or 8 sec under "Length".
- Tap "Create video". Generation runs for a few minutes, and Toybox refunds the credits if it fails.
- Watch the whole clip, then download the MP4. Results you don't save disappear after 24 hours.
A description that fits one 8-second clip:
A gray tabby cat stretches and yawns on a sunny kitchen windowsill, then settles back down with its paws tucked in. Soft early morning light, a potted herb beside the cat, shallow depth of field. Slow push-in at eye level that stops on the cat's face. Quiet kitchen sounds: a soft purr and a clock ticking on the wall. A single take with no words on screen.Every Video Creator clip comes with sound. With no separate sound setting, the dialogue, music or background noise you want goes into the description, written the way Google's prompt guide suggests. Listen before you share, since a spoken line isn't guaranteed to come out word for word. Video Creator makes 720p clips, currently costs 350 credits per clip and needs a Pro plan; see Toybox pricing. For 1080p, Video Creator Lite has a 1080p option for 8-second clips.