- AI APIs
- Releases
- Alerts
LLM model retirements in late 2026: shutdown dates and a migration plan for extension developers
OpenAI, Anthropic and Google retire LLM models from October to December 2026. The shutdown dates, the API changes that break swaps, and a plan for extensions.
- Author
- Extenify Team
- Published
- Reading time
- 14 min read

On this page
On September 30, 2026, Anthropic told developers that Claude Sonnet 4.5 will be retired on November 30. A day later, OpenAI deprecated GPT-5.1, GPT-5.3-Codex and GPT-5.4-Nano. And on October 23, OpenAI switches off a long list of legacy models, including gpt-4-turbo, gpt-4o-2024-05-13, o1, o3-mini and o4-mini. After a shutdown date, requests to those model IDs fail. They do not quietly fall back to something newer.
For a web app, a retired model is a config change and a deploy. For a browser extension that calls an LLM API, it can be a week of broken features: the model ID may be compiled into a package that has to pass store review, users pick up updates on their own schedule, and the first signal you get is often a one-star review that says "stopped working". This guide collects the shutdown dates that matter for the rest of 2026, the API changes that make "just swap the model name" fail, and a step-by-step plan to make the next retirement a non-event.
Key takeaways
- OpenAI's October 23 shutdown removes
gpt-3.5-turbo,gpt-4,gpt-4-turbo,gpt-4o-2024-05-13,gpt-4.1-nano,o1,o3-mini,o4-miniandgpt-image-1, among others. Claude Sonnet 4.5 follows on November 30, and older GPT-5 and o3 snapshots on December 11.- Swapping the model name is often not enough. Claude models from 4.7 onward return a 400 error for a non-default
temperature, and Google's Gemini 3.8 Flash migration guide says to striptemperature,top_pandtop_kand replacethinking_budgetwiththinking_level.- Keep the model ID out of the extension package. A remote configuration file is allowed under the Chrome Web Store Manifest V3 rules, and a server-side relay is better still.
- Treat a 404 for a model as an outage: fall back to the next model, record it, and tell bring-your-own-key users what to change.
- Watch your store reviews and rating in the days around each shutdown date. Retirement bugs show up there first.
The LLM shutdown calendar for the rest of 2026
These dates come from the official deprecation pages of OpenAI, Anthropic and the Gemini API, checked on October 3, 2026.
| Date | Provider | What goes away | Recommended replacement |
|---|---|---|---|
| Oct 2 (listed) | gemini-2.5-flash-image |
gemini-3.1-flash-image-preview |
|
| Oct 5 | antigravity-preview-05-2026 managed agent |
antigravity-preview-09-2026 |
|
| Oct 22 | veo-3.1-generate-preview, veo-3.1-fast-generate-preview, veo-3.1-lite-generate-preview |
gemini-omni-1.1-flash |
|
| Oct 23 | OpenAI | gpt-3.5-turbo, gpt-4, gpt-4-turbo, gpt-4-1106-preview, gpt-4o-2024-05-13, o1, o1-pro, o3-mini, o4-mini |
gpt-5.6-sol or gpt-5.6-terra |
| Oct 23 | OpenAI | gpt-4.1-nano, gpt-image-1, fine-tunes of gpt-3.5-turbo, gpt-4, gpt-4.1-nano, babbage-002 and davinci-002 |
gpt-5.6-luna, gpt-image-2.5-sunburst or -flare |
| Oct 31 | OpenAI | Existing Evals become read-only | Promptfoo, per OpenAI's guide |
| Nov 30 | Anthropic | claude-sonnet-4-5-20250929 |
claude-sonnet-5-5 |
| Nov 30 | OpenAI | Reusable prompt objects (v1/prompts), the Evals platform and Agent Builder |
Prompts in your own code, the Agents SDK |
| Dec 1 | OpenAI | gpt-image-1-mini, gpt-image-1.5, chatgpt-image-latest |
gpt-image-2.5-sunburst or -flare |
| Dec 11 | OpenAI | gpt-5-2025-08-07, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07, gpt-5-pro-2025-10-06, o3-2025-04-16, o3-pro-2025-06-10 |
gpt-5.6-sol, -terra or -luna |
Google notes that its dates are "the earliest possible dates" a model might be retired, so a listed Gemini date can slip, but you should not plan around that.
Two more things to put on the calendar even though they are not shutdowns:
- Gemini Flash introductory pricing ends December 31. Google's Gemini 3.8 Flash guide lists $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, then $1.50 and $7.50 from January 1, 2027. The same intro period covers 3.7 Flash and 3.6 Flash. If your free tier is priced on today's 3.8 Flash rates, that cost doubles on New Year's Day.
- Early 2027 is already booked. OpenAI removes
tts-1andtts-1-hdon January 6, legacy realtime and audio models on January 20,whisper-1and thegpt-4otranscription models on February 26, andgpt-5.1,gpt-5.3-codexandgpt-5.4-nanoon April 1, 2027.
Models to watch, not yet deprecated
Anthropic's model table lists Claude Haiku 4.5 as active with a retirement "not sooner than October 15, 2026", and Claude Opus 4.5 "not sooner than November 24, 2026". Anthropic promises at least 60 days' notice for publicly released models, so as long as no notice has gone out, neither can be retired before early December. If either one runs in your product, this is the quarter to test its successor.
Notice periods are not the same everywhere
How much warning you get depends on the provider and the kind of model:
- OpenAI gives at least 6 months for generally available models, at least 3 months for specialized variants such as
-chat-latestand Codex models, and as little as 2 weeks for anything withpreviewin the name. Its own advice is not to run business-critical production work on preview models "unless you can migrate on short notice." - Anthropic gives at least 60 days and emails customers with active usage of the model.
- Google announces deprecations in the Gemini API release notes and also limits some older models to existing users. Gemini 2.5 Pro and 2.5 Flash are not deprecated, but access is now limited "to users who have actively used them in the past", so a new project or a new API key may not be able to call them at all.
The emails go to whoever owns the API account. In a small team that is often a founder's inbox, not the person who maintains the extension.
Why a retired model hurts an extension more than a web app
A server-side app can change a model name and deploy in minutes. An extension has three extra delays between your fix and your users:
- Store review. The Chrome Web Store review process page says most reviews finish "within a few days, but it can take up to a few weeks." Edge Add-ons and Firefox Add-ons run their own queues.
- Update rollout. Chrome checks for extension updates "every few hours", per Chrome's extension hosting documentation, and a browser that stays closed does not check at all. For a while after any release your users are split across versions, which is covered in more detail in what to watch in the week after an extension release.
- Stored settings. Bring-your-own-key (BYOK) extensions usually save the selected model in
chrome.storage. Shipping a new default does not change the value a user already picked.
The result is a familiar pattern in issue trackers. In September, users of the Freelens AI extension reported that every request to Claude through an OpenAI-compatible gateway failed with a 400 "temperature is deprecated for this model", because the extension always sent temperature: 0 for any model that was not on its reasoning-model list. Roo Code users hit the same error "for every prompt in every scenario" when they selected claude-opus-4-7. And the Stack Overflow question about a 404 "models/gemini-1.5-flash is not found" has passed 45,000 views. The people asking are not doing anything unusual. Their code simply assumed a model ID and a parameter set would last.
Step 1: Find every model ID you ship
Start with an inventory. Search both your source and your built package, because SDKs and wrapper libraries can carry their own default model:
grep -rnoE "\b(claude-[a-z0-9.-]+|gpt-[a-z0-9.-]+|o[134](-mini|-pro)?|gemini-[a-z0-9.-]+|whisper-1|tts-1(-hd)?)\b" \
src/ dist/ | sort -u
Then check the other places a model ID hides:
- Default values for the settings page, and the list of models a BYOK user can pick from.
- Your relay or backend, if requests go through one, plus environment variables on each environment.
- Usage data. Anthropic's deprecation page describes an audit: export the Usage page in the Claude Console and review the CSV, which breaks usage down by API key and model. OpenAI and Google show per-model usage in their dashboards too.
Write the result down as a small table: model ID, where it is used, shutdown date, replacement, owner. That table is your migration plan.
Step 2: Move the model ID out of the package
The fix that makes every future retirement cheaper is to stop compiling model names into the extension.
Best option: a relay. If your extension already sends requests through your own server, which is also how you keep provider API keys out of the package (the same pattern as the relay in adding Google Analytics 4 to a Manifest V3 extension), the extension can ask for a capability such as "summarize" and let the server choose the model. A retirement then becomes a server deploy.
For BYOK extensions: a remote config file. When the extension calls the provider directly with the user's key, fetch the model settings as data. The Chrome Web Store Manifest V3 requirements explicitly allow "fetching a remote configuration file for A/B testing or determining enabled features, where all logic for the functionality is contained within the extension package." A JSON file with model names is data, not remotely hosted code.
// model-config.js
const CONFIG_URL = 'https://config.example.com/extension/models.json';
const CACHE_MS = 6 * 60 * 60 * 1000;
// Bundled fallback, used offline or if the remote file is invalid.
const BUNDLED = {
chat: {
model: 'claude-sonnet-5-5',
fallbacks: ['claude-sonnet-4-6'],
sendTemperature: false,
},
retired: {
'claude-sonnet-4-5-20250929': 'claude-sonnet-5-5',
'gpt-4o-2024-05-13': 'gpt-5.6-sol',
'o4-mini': 'gpt-5.6-terra',
},
};
const MODEL_ID = /^[a-z0-9][a-z0-9.\-]{1,63}$/;
function isValid(cfg) {
const chat = cfg?.chat;
return (
chat &&
MODEL_ID.test(chat.model) &&
Array.isArray(chat.fallbacks) &&
chat.fallbacks.every((m) => MODEL_ID.test(m)) &&
typeof chat.sendTemperature === 'boolean' &&
typeof cfg.retired === 'object'
);
}
export async function getModelConfig() {
const { modelConfig } = await chrome.storage.local.get('modelConfig');
if (modelConfig && Date.now() - modelConfig.fetchedAt < CACHE_MS) return modelConfig.value;
try {
const res = await fetch(CONFIG_URL, { cache: 'no-store' });
const value = await res.json();
if (!isValid(value)) throw new Error('invalid model config');
await chrome.storage.local.set({ modelConfig: { value, fetchedAt: Date.now() } });
return value;
} catch {
return modelConfig?.value ?? BUNDLED;
}
}
The retired map also fixes the stored-settings problem. When the user's saved model is on it, swap it for the replacement and tell them once:
// background.js
import { getModelConfig } from './model-config.js';
async function migrateSavedModel() {
const { retired } = await getModelConfig();
const { model } = await chrome.storage.sync.get('model');
const replacement = model && retired[model];
if (!replacement) return;
await chrome.storage.sync.set({ model: replacement, modelNotice: { from: model, to: replacement } });
}
chrome.runtime.onStartup.addListener(migrateSavedModel);
chrome.runtime.onInstalled.addListener(migrateSavedModel);
Because the map is fetched, you can add next month's retirements without shipping a new version.
Step 3: Stop sending parameters new models reject
Most migration failures in 2026 are not the model name. They are the request body.
- Anthropic. The deprecations page lists
temperature,top_pandtop_kas deprecated for Claude Opus 4.7 and later. On those models a non-default value returns a 400 error, and the replacement for Sonnet 4.5, Claude Sonnet 5.5, is one of them. The Python SDK from version 1.0 removes the three parameters entirely, so passing them raises aTypeErrorbefore any request is sent. - Google. The Gemini 3.8 Flash migration checklist says to strip
temperature,top_pandtop_k, replacethinking_budgetwith the stringthinking_level(low,mediumorhigh;minimalis not supported on 3.8 Flash) and removecandidate_count. A Hacker News thread from July quotes Google's guidance at the time: in future model generations, "supplying these parameters returns an HTTP 400 error."
The robust pattern is to stop sending sampling parameters by default, and only add them when the model's config says it accepts them. Matching on model name patterns, which is what broke the Freelens extension, fails the next time a provider ships a model your pattern did not expect.
function buildClaudeRequest(chat, model, messages) {
const body = { model, max_tokens: 1024, messages };
// Only send temperature when the config for this model allows it.
if (chat.sendTemperature) body.temperature = 0.2;
return body;
}
If you need steadier output, both Anthropic and Google now point to the prompt and system instructions as the place to ask for it.
Step 4: Handle "model not found" as an outage
Even with remote config, a request will eventually hit a model that no longer exists: a cached config, a user on an old build, or a date you missed. Make that path explicit.
Each provider reports it differently. Anthropic returns a 404 with an error type of not_found_error, raised as NotFoundError by its SDKs. OpenAI's SDKs raise NotFoundError for a resource that does not exist. Gemini returns a 404 NOT_FOUND with a message like "models/gemini-1.5-flash is not found". Here is a relay-side handler for Claude that tries the fallbacks in order:
// relay: complete.js (runs on your server, holds the API key)
export async function complete(chat, messages, env, log) {
for (const model of [chat.model, ...chat.fallbacks]) {
const res = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
headers: {
'content-type': 'application/json',
'x-api-key': env.ANTHROPIC_API_KEY,
'anthropic-version': '2023-06-01',
},
body: JSON.stringify(buildClaudeRequest(chat, model, messages)),
});
if (res.ok) return res.json();
const err = await res.json().catch(() => ({}));
if (res.status === 404 && err?.error?.type === 'not_found_error') {
log.warn('model_unavailable', { model });
continue; // try the next model
}
throw new Error(`LLM request failed: ${res.status} ${err?.error?.type ?? ''}`);
}
throw new Error('No configured model is available');
}
Two details matter. First, log the event with the model name, so a missed retirement shows up as an alert in your logs instead of a pile of support emails. Second, for BYOK users, turn the error into a sentence they can act on, such as "The model you selected was retired by the provider. Choose a newer model in Settings." A raw 404 in a popup reads as "this extension is broken."
Step 5: Test the replacement on your own prompts first
A new model that answers the same request is not the same as a new model that does your job as well. A Hacker News commenter put it plainly in August: "It's almost never a drop-in replacement, and having to check and adjust integrations and workflows with new models gets old fast." Another, discussing the Gemini 3.6 release, warned that "highly specific workflows may rely on specific invisible aspects of a model."
A short check before each switch catches most of it:
- Collect 20 to 50 real inputs your extension handles: a mix of short, long, non-English and awkward pages. Strip anything personal before you store them.
- Run them through the current and the replacement model with the same prompt, and compare the outputs side by side. Check format first (does your parser still work?), then quality.
- Compare token usage and latency. Google's guide notes that Gemini 3.8 Flash "can use more tokens on longer running and complex tasks, by design", which changes your cost per request even at a lower price per token.
- Ship the new model to a slice of users through the remote config, then to everyone.
If the extension sends page content to an LLM at all, check that your store disclosures still match what you send. The rules on what you must disclose are covered in Chrome Web Store changes in 2026.
Step 6: Watch the stores around each shutdown date
Users of a broken AI feature rarely file a bug. They leave a review. That makes your store listings the most reliable early warning you have, especially for BYOK users you cannot see in your own logs.
In the days around each date in the calendar above, watch for:
- New one- and two-star reviews that mention "stopped working", "error", "not responding" or a model name.
- An unusual burst of reviews on one store, which can point at a build or setting that only that store's users have.
- A rating drop on an established listing. It arrives last and recovers slowest, as explained in reading review and rating trends.
Extenify sends these as daily change alerts for every store your extension is listed on, to email, Slack, Discord, Telegram or Microsoft Teams. Setting up store change alerts your team will actually read shows which ones to route to the person who owns the AI integration.
A migration checklist for the next 90 days
- List every model ID in source, build output, settings defaults and backends.
- Map each one to its shutdown date and replacement from the official pages.
- Move model selection into a relay or a validated remote config file.
- Add a
retiredmap that migrates saved user settings. - Stop sending
temperature,top_pandtop_kunless the model config allows them. - Replace Gemini
thinking_budgetwiththinking_levelbefore moving to 3.8 Flash. - Handle a 404 for a model with a fallback, a log event and a readable message.
- Run your prompt set against each replacement, including token cost.
- Recheck pricing for January 1, 2027.
- Turn on review and rating alerts before October 23, November 30 and December 11.
Frequently asked questions
What happens when an LLM model is retired?
Requests that use the retired model ID fail. Anthropic and OpenAI both say the model is no longer available after the shutdown date, and Google says the endpoint is "completely turned off". There is no automatic redirect to the replacement on these three platforms, so your code has to change the model ID itself.
Which OpenAI models shut down on October 23, 2026?
According to OpenAI's deprecations page: gpt-3.5-turbo-0125, gpt-4-0613, gpt-4-1106-preview, gpt-4-turbo, gpt-4.1-nano, gpt-4o-2024-05-13, gpt-image-1, o1, o1-pro, o3-mini and o4-mini, along with their aliases and several fine-tuned models. The recommended replacements are gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna, and the GPT Image 2.5 models.
When is Claude Sonnet 4.5 retired?
Anthropic announced on September 30, 2026 that claude-sonnet-4-5-20250929 will be retired on the Claude API on November 30, 2026, with claude-sonnet-5-5 as the recommended replacement. Amazon Bedrock and Google Cloud set their own schedules, so check those separately if you use them.
Why do I get "temperature is deprecated for this model"?
Claude models from 4.7 onward return a 400 error when temperature, top_p or top_k is set to a non-default value. Remove the parameter from requests to those models, and only send it when your configuration says the model accepts it.
Can a Chrome extension change its model without a store update?
Yes, if the model name is data. The Chrome Web Store allows extensions to fetch a remote configuration file as long as all the logic stays in the package. A JSON file that names the model, its fallbacks and which parameters to send is configuration, not remote code.
Summary
The last quarter of 2026 is full of model shutdowns: OpenAI's legacy cleanup on October 23, Claude Sonnet 4.5 on November 30, image models on December 1, the original GPT-5 and o3 snapshots on December 11, and a Gemini price change on January 1. None of them should break your extension if the model ID lives outside the package, requests only send parameters the model accepts, a missing model triggers a fallback instead of an error message, and you test each replacement on your own prompts. The last line of defense is the one your users write: watch your reviews and ratings around each date, and you will hear about a missed migration in a day instead of a week.
Image credits
- Cover photo: Wooden hourglass with pink sand beside a yellow pen holder by Rashed Hossain on the WordPress Photo Directory, CC0.


