AI model by Google
Veo 3.1 for business
Google's video generation model, creating short video with native audio from text and images. Available as a preview, with fast and lite versions.
- Reviewed on September 24, 2026
Capability tiersRelative, not benchmarks
Key facts about Veo 3.1
- Model ID at review
- veo-3.1-generate-preview
- Provider
- Main category
- Image and video
- Open weights
- No, available as a hosted service
- Last reviewed
- September 24, 2026
Fit
Where it fits and where it does not
Good at
-
Short video clips from text prompts.
-
Animating a still image.
-
Video with native audio.
-
Extending existing clips.
-
Access through the Gemini API.
Not the right choice for
-
Long-form video without editing.
-
Final brand content without review.
-
Realistic video of real people without consent.
-
Low budgets at high volume, since video is priced per second.
Use cases
Business use cases we would use it for
Social clips
Storyboards
Training content
Our notes
When we would choose it
Veo 3.1 is Google's current video generation model in the Gemini API. It creates video with native audio from text and image input, and supports video extension and generation from specific frames. Google lists a standard, a fast and a lite version, all as previews at the time of review. Veo 2.0 and 3.0 were shut down on June 30, 2026.
Video generation has become good enough for short social clips, storyboards and simple explainer scenes. It is not yet a replacement for a planned shoot or an editor. The cost is per second of video, so volume adds up quickly, and every clip needs review before it is used.
We would use Veo inside a workflow with a clear brief, prompt templates that match your brand, a review queue and a record of what was generated. Human approval stays mandatory for anything that reaches customers.
For comparison, Runway offers video generation and video editing models with its own API. For still images, see Nano Banana and GPT Image. See content generation.
Keep in mind that preview models can change behavior between updates. For a campaign that runs over weeks, generate and approve the clips early, and store the final files rather than relying on generating the same result again.
Before you commit
Things to check before you commit
-
Preview status
Veo 3.1 is offered as a preview. Terms and behavior may change.
-
Cost per second
Video is priced per second of output. Estimate carefully.
-
Usage terms
Check commercial use terms and content policies.
-
Older versions
Veo 2.0 and 3.0 shut down on June 30, 2026.
Alternatives
Models to compare it with
Runway Gen-4.5
Runway's video models: Gen-4.5 for text-to-video and image-to-video, and Aleph 2.0 for editing existing video, with a developer API.
Nano Banana (Gemini image)
Google's Gemini-native image generation and editing models, known as Nano Banana 2 and Nano Banana Pro, which replace Imagen in the Gemini API.
GPT Image 2.5
OpenAI's current image generation and editing models: Sunburst for the highest quality and Flare for fast everyday images.
Keep exploring
Solutions, services and guides
Related services
View all related services- AI workflows Step-by-step automations where AI handles the reading, sorting and drafting inside a process you control.
- WordPress development Custom WordPress themes, plugins and cleanups for teams who want to keep editing content themselves.
- Computer vision Software that reads photos and scans: damage checks, stock counts, document capture and quality control.
Solutions
View all solutions- AI-assisted content workflows Briefs, first drafts, edits and repurposing in a workflow where people set the angle and approve every word.
- Scheduling and drafting automation Turn one piece of content into posts for each channel, schedule them and track results, with approval before anything goes out.
- AI document processing Read forms, applications, IDs and statements, extract the fields you need and route each document for the right review.
- AI translation workflows Translate websites, products and support content quickly with AI, a shared glossary and native-speaker review where it matters.
Industries
View all industriesCase studies
View all case studiesMCP servers
View all mcp serversGlossary terms
View all glossary terms- Multimodal model A multimodal model is an AI model that can take in more than one kind of input, such as text with images, audio or video.
- Prompt engineering Prompt engineering is the practice of designing, testing and improving prompts so AI models produce reliable results for a specific task.
- Large language model A large language model, or LLM, is an AI model trained on vast amounts of text that can understand and generate language, and often images and code.
- Prompt A prompt is the instruction and context you give an AI model to get the output you want.
FAQ
Questions people ask us
Have a question that is not here? Ask us directly.
Yes. Google lists native audio with the generated video.
At review time it was offered as a preview in the Gemini API.
By the second of generated video. Check the Gemini pricing page for current rates.