Gemini API Preview Shutdowns: Veo 3.1 Moves to Omni, Voice Follows in November

Veo 3.1 and Omni previews may end from October 22, TTS and Live previews from November 17. Cost per clip, voice replacements and the January price rise.

Google’s Gemini API deprecations page, updated October 7, lists four video preview model IDs with an earliest shutdown date of October 22, 2026 and five voice preview IDs with an earliest date of November 17. Gemini API deprecations The page is explicit that these are floors, not fixed dates: “The shutdown dates listed in the table indicate the earliest possible dates on which a model might be retired. We will communicate the exact shutdown date to users with advance notice.” So every date below reads “no earlier than”. Plan for the floor, because Google has not promised anything later.

The three Veo 3.1 preview endpoints are replaced by gemini-omni-1.1-flash, a different model, served through the Interactions API, that bills per output token instead of per second, so the cost of a clip changes shape. The replacement text-to-speech models carry a scheduled price doubling on January 1, 2027, so a voice pipeline that migrates in November gets one price for six weeks and another after that. Gemini API pricing

S5 Labs has not run these models. Every price below is from Google’s pricing page as of October 7, and every capability claim is from Google’s docs.

The calendar

Earliest shutdownModel IDReplacementWhat changes
October 5 (done)antigravity-preview-05-2026antigravity-preview-09-2026Built-in tool names and parameters changed; see the September 17 release note
No earlier than October 22veo-3.1-generate-preview, veo-3.1-fast-generate-preview, veo-3.1-lite-generate-previewgemini-omni-1.1-flashNew model, Interactions API, per-token billing
No earlier than October 22gemini-omni-flash-previewgemini-omni-1.1-flashSame price; deprecated September 30
No earlier than November 17gemini-3.1-flash-tts-preview, gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-ttsgemini-3.8-flash-tts or gemini-3.8-flash-lite-ttsLower audio price until December 31, then double
No earlier than November 17gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025gemini-3.8-liveSame audio prices; text prices rise; async function calling by default
January 1, 2027Gemini 3.8 Flash TTS and Flash-Lite TTS(price change, no shutdown)Input, audio output and caching prices double

Sources: Gemini API deprecations, Gemini API release notes, Gemini API pricing.

The Antigravity row is already past. Its release note says a remote-sandbox caller that only reads output_text changes the agent string and nothing else; anyone running tools locally or parsing function_call steps has renamed tools, PascalCase parameters and line-range file edits to handle. Gemini API release notes

The Veo previews end on the Gemini API; Veo continues on Vertex

The deprecations table pairs each Veo 3.1 preview with gemini-omni-1.1-flash, which went GA on August 27. Gemini API deprecations Gemini API release notes That is a replacement model, not a GA version of Veo. Google’s own video landing page, last updated June 30, still says to “Use Gemini Omni Flash as your default model for video generation” and to use Veo 3.1 “for specific capabilities like scene extension, last-frame control, or integration with legacy pipelines”. Video generation in the Gemini API Nothing on the Veo 3.1 guide or model page mentions October 22; the only notice is the deprecations table.

Veo continues elsewhere. The Gemini Enterprise Agent Platform (Vertex) lists veo-3.1-generate-001 and veo-3.1-fast-generate-001, released November 17, 2025, with a retirement date of “November 17, 2026 or later” and no replacement named; the page adds that timelines may be extended but will not move earlier. Model versions and retirement dates The Gemini API’s own table pointed Veo 3.0 users to “the GA models on the Gemini Enterprise Agent Platform” when those shut down in June; it does not name the IDs, so treat the Vertex IDs above as the likely targets and confirm in the console. So a team that needs Veo specifically has a path, but it is a different platform with its own billing, and the Vertex IDs carry their own one-year floor.

The capability lists differ as much as the prices do. Veo 3.1 generates clips of up to 8 seconds at 720p, 1080p or 4K (no 4K on Lite), extends previously generated Veo videos at 720p only, takes first and last frames, and accepts up to three reference images. Lite also lacks extension and reference images. Veo 3.1 guide Omni 1.1 Flash runs through the Interactions API. Its preview launch note describes 3 to 10 second videos at 720p, and the guide says generated videos can be extended by 10 seconds at a time up to 40 seconds. Gemini API release notes Gemini Omni Flash guide Resolution is 360p, 720p (default), 1080p or 4K, but the August 27 note says 1080p and 4K “are generated using upscaling”. The guide’s limitations list is the part to read before committing: editing or extending uploaded videos is unavailable in the EEA, Switzerland and the UK; uploaded videos for editing must be 10 seconds or less; video references are capped at three clips of three seconds each, with their audio ignored; and system instructions, temperature, top_p, stop sequences and negative prompts are not supported. Gemini Omni Flash guide

On audio input, Google’s pages pull in two directions. The video landing page and the Omni guide both describe audio as a native input, and the pricing table lists audio input at the same $1.50 per million tokens as text, image and video. Gemini Omni Flash guide Gemini API pricing The same guide’s limitations say “Uploading audio references is unsupported in the current version of the API” and “Voice editing is not supported.” Read together: audio is a priced, documented modality, and the current API does not let you upload an audio reference to drive a generation. If your Veo workflow never sent audio in, this changes nothing. If you were planning to, test it before you depend on it.

Cost per clip: per second against per token

Veo 3.1 bills per second of generated video, with audio included by default, and charges only when a video is successfully generated. Gemini API pricing Omni bills on output tokens: “5,792 tokens per second of 720p video”, at $17.50 per million video output tokens, which Google rounds to “approximately $0.10 per second”. The exact figure is 5,792 × 17.50 / 1,000,000 = $0.10136 per second. The pricing page gives the token rate for 720p only, so the Omni column below stops there; 1080p and 4K Omni output are upscaled and have no published token rate.

The clip length is 8 seconds, the longest Veo 3.1 offers. Input tokens are left out because a text prompt of a few hundred tokens at $1.50 per million is a fraction of a cent.

720p, 8-second clip, audio includedBasisPer secondPer clipPer 1,000 clips
Veo 3.1 Standard (preview)$0.40/s$0.40$3.20$3,200
Veo 3.1 Fast (preview)$0.10/s$0.10$0.80$800
Veo 3.1 Lite (preview)$0.05/s$0.05$0.40$400
Gemini Omni 1.1 Flash5,792 tokens/s × $17.50/1M$0.1014$0.81$811

Prices from Gemini API pricing; clip arithmetic is ours.

So at 720p the move is close to neutral for Fast users ($0.80 to $0.81), a 75% cut for Standard users ($3.20 to $0.81) and a doubling for Lite users ($0.40 to $0.81). Lite is the tier that loses: it launched March 31 as Google’s “most cost-efficient video generation model, designed for rapid iteration and building high-volume applications”, and the replacement costs twice as much per second with no cheaper Omni tier on the page. Gemini API release notes

For the higher resolutions, only Veo has a published figure:

Veo 3.1 tier, 8-second clip720p1080p4K
Standard$3.20$3.20$4.80
Fast$0.80$0.96$2.40
Lite$0.40$0.64not offered

A 4K Standard clip at $4.80 has no stated Omni equivalent. Omni’s 4K is an upscale of a 720p generation, and whether it bills at the 720p token rate or something higher is not on the pricing page. Ask Google, or run one and read the usage metadata, before quoting a client.

Two other cost differences do not show in a per-clip table. Omni’s headline feature is conversational editing, but the pricing page says nothing about how edits are metered beyond “total output token consumption”, so read the usage metadata on a few edit turns before you budget on the per-clip figure. And a 40-second Omni clip built from one 10-second generation and three 10-second extensions produces about 231,680 output tokens, around $4.05 at 720p, plus whatever the input side of each extension costs; the page gives no video-input token rate for Omni, so that part is unpriced.

The Omni preview itself is the easy row: gemini-omni-flash-preview has identical prices to the GA model, so moving to gemini-omni-1.1-flash is an ID change and a re-test. Gemini API pricing

Voice: TTS gets cheaper now and dearer in January

Three TTS previews reach their floor on November 17, all pointing at gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts, which went GA on September 22 with a Voices endpoint, voice design and consent-verified voice replication. Gemini API deprecations Gemini API release notes Google’s TTS guide names Flash-Lite as the replacement for gemini-3.1-flash-tts-preview, and the release note positions it “for high-throughput production and real-time voice agent cascades”; Flash TTS is positioned for “studio-grade voice fidelity, nuanced acting, regional dialects, and long-form multi-turn stability”. Gemini API release notes

Google bills audio at 25 tokens per second, so an hour of generated speech is 90,000 output tokens. Gemini API pricing The table converts each model’s standard-tier audio output price to that hour. Text input is listed separately because it is small: a 10,000-token script costs half a cent at $0.50 per million.

ModelStatusText input / 1MAudio output / 1MPer hour of audio
gemini-2.5-flash-preview-ttsshutdown no earlier than Nov 17$0.50$10.00$0.90
gemini-2.5-pro-preview-ttsshutdown no earlier than Nov 17$1.00$20.00$1.80
gemini-3.1-flash-tts-previewshutdown no earlier than Nov 17$1.00$20.00$1.80
gemini-3.8-flash-ttsGA; prices through Dec 31$0.50$9.00$0.81
gemini-3.8-flash-ttsfrom Jan 1, 2027$1.00$18.00$1.62
gemini-3.8-flash-lite-ttsGA; prices through Dec 31$0.50$6.00$0.54
gemini-3.8-flash-lite-ttsfrom Jan 1, 2027$1.00$12.00$1.08

Prices from Gemini API pricing; per-hour figures are ours, matching Google’s own “per 10s audio” equivalents ($0.00225 for Flash TTS through December 31, $0.0045 after).

The direction depends on where you start. From the 3.1 preview or 2.5 Pro preview at $1.80 an hour, Flash TTS is a 55% cut until December 31 and still 10% cheaper from January. From the 2.5 Flash preview at $0.90, Flash TTS is a small cut now and an 80% rise from January; Flash-Lite is a 40% cut now and a 20% rise from January. Batch and Flex tiers are half the standard price in both periods, and context caching doubles on the same date. The pricing page does not use the word “introductory”, but the structure is the same one Google used for Gemini 3.8 Flash, where a launch price ends in December.

There is code to change as well. The TTS guide’s current request shape passes the transcript as input, styling as a speech_metadata annotation and the voice in generation_config.speech_config; its examples go through the Interactions API, and it asks for google-genai 2.25.0 or later, or REST. Text-to-speech guide Multi-speaker output in one request is limited to two prebuilt voices; custom or replicated voices in a dialogue have to be synthesized turn by turn and concatenated as raw PCM. Stored custom voices are capped at 200 per project and expire one year after last use.

Live: same audio price, new defaults

The two Live previews, gemini-3.1-flash-live-preview and gemini-2.5-flash-native-audio-preview-12-2025, point at gemini-3.8-live, GA since September 15 alongside a gemini-3.8-live-extended-thinking variant for background reasoning during a call. Gemini API deprecations Gemini API release notes

Audio prices do not move. Google prices 3.8 Live and the 3.1 preview on one table: $3.00 per million audio input tokens (or $0.005 per minute), $12.00 per million audio output ($0.018 per minute), $0.75 text in, $4.50 text out and $1.00 per million for image or video input ($0.002 per minute). The 2.5 native-audio preview has the same $3.00 and $12.00 audio rates but cheaper text: $0.50 in and $2.00 out. Gemini API pricing A voice agent that streams audio both ways pays the same per minute after the move; one that pushes a lot of text context or reads a lot of text output pays more, with text output more than doubling for 2.5 users.

The behavior changes are the ones to test. The 3.8 Live release note lists “interleaved reasoning, default asynchronous function calling, and full session client content updates”. Gemini API release notes Function calling that was synchronous on a preview model is asynchronous by default on the replacement, which changes how a tool result lands mid-conversation. Our GPT-Live-1 piece made the same point about OpenAI’s voice stack: the price per minute is the easy comparison, and the delegation model is where the integration work sits.

A checklist with dates

  1. Now. Grep every codebase, config and prompt store for the nine preview IDs in the calendar table. Include gemini-omni-flash-preview, which already costs the same as GA and is the quickest fix.
  2. Before October 22. Move Veo 3.1 preview calls to gemini-omni-1.1-flash on the Interactions API, or to the Veo 3.1 IDs on the Gemini Enterprise Agent Platform if you need Veo’s extension, last-frame control or 4K pricing. Re-run a sample of real prompts and read the output token counts; the $0.81 per 8-second clip is Google’s rate, not a measurement of your prompts.
  3. Before October 22, Lite users. Price the doubling against Omni’s edit loop. If a job needed three Lite generations to get one keeper, one Omni generation plus an edit may cost about the same; if Lite one-shotted it, the bill rises.
  4. Before November 17. Swap TTS previews to 3.8 Flash or Flash-Lite TTS through the new request shape, then regenerate your reference clips and listen; the voices come from a different model. Decide whether to use the Batch tier for anything not interactive.
  5. Before November 17. Move Live previews to gemini-3.8-live and test every tool call under asynchronous function calling. Compare text-token usage per session, since that is the price that moved.
  6. Before January 1, 2027. Re-forecast TTS spend at $18 and $12 per million audio tokens. A narration product budgeted on $0.81 an hour in November costs $1.62 an hour in January.
  7. Keep checking the page. Google says it will announce exact dates with notice; the deprecations page is where those land, and it was updated on October 7.

The same week brings the OpenAI and GitHub Copilot retirement dates, so a team that runs both vendors has two checklists with overlapping deadlines.

Key Details

ItemDetail
Source pageGemini API deprecations, last updated October 7, 2026; dates are “earliest possible”
Video previews, no earlier than Oct 22veo-3.1-generate-preview, veo-3.1-fast-generate-preview, veo-3.1-lite-generate-preview, gemini-omni-flash-preview → gemini-omni-1.1-flash (GA Aug 27)
Veo 3.1 pricePer second, audio included: Standard $0.40 (720p, 1080p) / $0.60 (4K); Fast $0.10 / $0.12 / $0.30; Lite $0.05 / $0.08, no 4K; charged only on success
Omni 1.1 Flash price$1.50 per 1M input (text, image, video, audio); $9.00 text out; $17.50 video out at 5,792 tokens per second of 720p, about $0.10/s; no token rate for 1080p or 4K
8-second 720p clipVeo Standard $3.20, Fast $0.80, Lite $0.40; Omni about $0.81
TTS previews, no earlier than Nov 17gemini-3.1-flash-tts-preview, gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-tts → gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts (GA Sept 22)
Gemini 3.8 TTS priceFlash: $0.50 in / $9.00 audio out through Dec 31, $1.00 / $18.00 from Jan 1, 2027. Flash-Lite: $0.50 / $6.00, then $1.00 / $12.00. 25 audio tokens per second
Live previews, no earlier than Nov 17gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025 → gemini-3.8-live (GA Sept 15)
Gemini 3.8 Live price$3.00 audio in ($0.005/min), $12.00 audio out ($0.018/min), $0.75 text in, $4.50 text out, $1.00 image/video in
Already shut downantigravity-preview-05-2026 on October 5 → antigravity-preview-09-2026
Veo on Vertexveo-3.1-generate-001 and veo-3.1-fast-generate-001, released Nov 17, 2025, retirement “November 17, 2026 or later”

Sources

Continue reading.

Insight13 min read

Anthropic's Cyber Verification Program Now Has Three Tiers. Which One to Apply For

Anthropic split its Cyber Verification Program into Defense, Red Team and Specialized tiers. Who qualifies, review times, the retention catch and Bedrock.

Insight14 min read

Claude Haiku 5.5: A 90% List-Price Cut, a 75% Real One, and Five Ways to Get a 400

Haiku 5.5 lists at a tenth of Haiku 4.5's price; Anthropic says about 75% less in practice. The gap, the 400s after switching, and Sonnet 5.5 vs Luna.

Insight14 min read

Mistral Large 4: What Is Confirmed Before the Weights Drop, and What Is Not

Mistral Large 4 is a 1T-parameter MoE you can call today at a half-price preview. Weights, license and independent scores are still due. What to check.