An AI video with sound has audio made by the same model that draws the picture: footsteps that land when the feet do, a spoken line that matches the lips, rain you can hear on the roof. As of September 2026, Google's video models do this natively, and you steer the sound with words in the prompt. Toybox AI's Video Creator clips come with sound as well. The sections below cover which tools generate audio, how to write dialogue, effects and ambience as cues, and when to add your own track in an editor.
Which tools make AI video with sound
Google documents native audio for its current video models, and several Google apps give you access to them. Pick an AI video with sound maker by where you'll post and how much you want to spend. As of September 2026:
| Tool | Sound | Length of one clip | Cost to start |
|---|---|---|---|
| Google Flow | Runs Veo 3.1 and Gemini Omni Flash models; Google documents native audio for both | 4, 6 or 8 seconds (Veo 3.1); up to 10 seconds (Gemini Omni Flash) | 50 free Flow credits a day (18 and older); a Veo 3.1 Lite clip costs 10 |
| Gemini app | Gemini Omni, which Google lists with native audio generation | 10 seconds | A Google AI plan, from $4.99 a month in the US (Google AI Plus); 18 or older |
| Gemini API | Veo 3.1 models, with audio always on | 4, 6 or 8 seconds | Paid, for developers |
| YouTube app | "Create video" for Shorts; YouTube suggests describing any audio you need | Up to 10 seconds | Built into the app; not in the European Economic Area |
| Google Photos | Photo to video clips may now contain audio | 6 seconds | A daily limit, which a Google AI plan raises |
| Toybox Video Creator | Clips come with sound, shaped by your prompt; no separate sound setting | 4, 6 or 8 seconds | A Pro plan; currently 350 credits per clip |
OpenAI's Sora is no longer an option: the Sora app and website closed on April 26, 2026, and OpenAI pulled its Sora video models from the API on September 24, 2026. How the two model families differed is covered in Sora vs Veo.
How to prompt dialogue, sound effects and ambience
Describe the sound in the same prompt as the picture, using the three kinds of cue Google's Veo 3.1 guides name: dialogue, sound effects and ambient noise. Google's API documentation says Veo turns these cues into a soundtrack that is synced to the picture.
- Dialogue goes in quotation marks. Name who speaks and how: The baker says quietly, "First batch is always the best."
- Sound effects get a label. Write "SFX:" and one clear sound tied to an action, such as "SFX: the cap twists off with a sharp hiss."
- Ambience sets the room. Write "Ambient noise:" and the background you want: an oven fan humming, traffic two floors down, waves on a pier.
- Music is a style, not a song. Give the genre, the tempo and one or two instruments, such as "slow acoustic guitar and soft piano". Never name a real song or artist.
- Timing can be spelled out. Google's Veo 3.1 guide shows timestamp prompting, with bracketed segments such as [00:00-00:02] for each shot and its sounds.
These prompts follow the cue format in Google's Veo 3.1 guides, so they suit Veo 3.1 in Flow or the Gemini API. Each one also runs in Toybox's Video Creator, which is powered by Google Veo 3.1. Its clips come with sound, and there's no separate setting for it: the cues in your prompt are how you ask for voices, effects and music. Spoken lines may not match your wording exactly, so play each downloaded clip with the sound on before you post it. More ready-made scenes are in the collection of Veo 3 prompts.
Two people talking
Medium two-shot in a small bakery at dawn. The baker, a woman in a flour-dusted apron, slides a tray of croissants onto the counter and says, "First batch is always the best." Her young assistant laughs and replies, "Then I'm taking two." SFX: the metal tray scrapes the counter. Ambient noise: an oven fan hums. Warm morning light.Sound effects that match the action
Close-up of a glass bottle of sparkling lemonade on a picnic table. A hand twists the cap off. SFX: a sharp fizzing hiss, then bubbles crackling as the lemonade is poured over ice in a tall glass. Ambient noise: birds and a light breeze in the trees. Bright summer afternoon, shallow depth of field.Ambience with no voices
Wide shot of a quiet mountain lake at sunrise, mist drifting over the water, a wooden canoe tied to a dock. Ambient noise: loons calling across the lake, water lapping against the dock, wind moving through the pines. Only these natural sounds, with no music and no voices. Slow pan from left to right.A timed sequence with a music cue
[00:00-00:03] Close-up of a runner lacing her shoes on a wet city sidewalk before dawn. SFX: laces pulled tight. [00:03-00:06] Tracking shot as she starts running past glowing shop windows. SFX: steady footsteps on wet pavement. [00:06-00:08] Wide shot from behind as she turns onto a bridge, and an upbeat electronic track with a steady beat begins.Tips for better sound in AI video
- Keep spoken lines short. An 8-second clip holds one or two short lines. Read yours out loud with a timer, and cut it if it runs past about six seconds, so the speaker isn't cut off.
- Write speech in English. Google's Veo documentation for the Gemini API lists English as fully supported and other languages as not yet evaluated, so results in them can vary.
- Tie every sound to something on screen. "SFX: the door slams as she leaves" gives you a moment to check. A sound with no source in the picture is hard to judge and easy to get wrong.
- Describe the quiet you want. Google's Veo 3.1 guide advises writing an exclusion into a concrete description of the scene instead of a vague ban. For sound, name what should be heard and rule out the rest: "only wind and gulls, with no music and no voices".
- Make each speaker easy to tell apart. When two characters talk, give each a clear description ("the older man in the gray coat"), so you can see whose line is whose when you check the lip sync.
- Listen twice. Play the clip on headphones, then on a phone speaker. Mumbled words and clipped endings are easiest to catch on the phone.
- Expect some failed tries. Google says Veo 3.1 sometimes blocks a video because of safety filters or problems with the audio, and the Gemini API doesn't charge for a blocked video. In Toybox, a failed generation refunds its credits automatically.
For how to describe music and mood across a whole project, the Veo prompting guide goes further.
Limits to know before you rely on AI sound
- The prompt is the only sound control in Toybox. Video Creator (powered by Google Veo 3.1) and Video Creator Lite (Veo 3.1 Lite) clips come with sound, but there's no volume, mute or music setting. To change what you hear, change the sound lines and run the prompt again.
- Spoken lines may not be word-perfect. A word can change or drop out, especially in a long line. Listen to every line before you post, and shorten any that come out wrong.
- A clip is 4, 6 or 8 seconds. Dialogue has to fit inside it, and adding an end frame makes a Toybox clip 8 seconds.
- Voices are the model's, not yours. Toybox has no voice upload. Google Flow has preset voices and custom voices built from a preset, and the Gemini app can use your own voice through its avatar feature.
- Realistic AI audio may need a label. YouTube asks for its "AI use" disclosure on AI content that looks or sounds real, and its examples name AI-made music.
- Both video tools are Pro-only. A Video Creator clip currently costs 350 credits; a Video Creator Lite clip costs 180 at 720p or 220 at 1080p. See current prices.
When to add your own soundtrack
Add your own track in a video editor when you want your own voice, a full song, or one soundtrack running across several clips. Any editor that has a timeline with separate audio tracks works, whether it runs on a phone or a laptop.
- A voiceover. Record it on your phone in a quiet room, close to the mic. Write it first and time it against the clip, so each line lands on the right shot.
- Music from a library. YouTube Studio's Audio Library has music that YouTube says is free to use in Shorts and won't receive a copyright claim. Read the license terms of any other library before you use a track.
- Music you generate. Music Generator turns a song title and a short description into two MP3 versions. Choose "Instrumental" for background music. It's a Pro tool, currently 95 credits a run, and there's no length setting, so cut the track to fit in your editor.
- Sound effects you record. Record the real thing when you can, such as a door, a pour or footsteps, and line it up with the frame where the action happens.
- Sound effects you generate. Text-to-sound tools make a single effect from a description. In ElevenLabs' Sound Effects tool, for example, you type something like "footsteps on wet gravel", set a length of up to 30 seconds or leave it on auto, and get four versions to download as MP3 or WAV.
Keep music lower than any voice, and fade it out on the last second so the clip doesn't end on a cut-off note. If the clip's own sound competes with your track, lower or mute the clip's audio in the editor.
How to make a Video Creator clip with sound
- Open Video Creator from a Pro account. The video tools aren't part of the Free plan.
- Write the scene into "Describe your video". Cover the camera, subject, action and setting in up to 2,000 characters, and add a starting image through "Start from a photo" if you have one.
- Add the sound lines. Put spoken words in quotes after the speaker, start effects with "SFX:" and the background with "Ambient noise:". Name the music by genre, tempo and instrument.
- Size the clip for the dialogue. Pick Wide (16:9) or Vertical (9:16), and choose 8 sec when a spoken line needs room.
- Press "Create video". Allow a few minutes. If the clip fails, the credits are returned automatically.
- Download the MP4 and play it with the sound on. Check every spoken word, and check that each sound lands on its action.
- Fix what's off. Shorten a line that comes out wrong and run the prompt again. For your own voiceover or a Music Generator track, lower the clip's audio in your editor and lay the new track under it.