- AI APIs
- Manifest V3
- Analytics
How to stream AI responses in a Chrome extension: SSE, ports and the Manifest V3 service worker
How to stream AI responses in a Chrome extension. Relay SSE from your backend over a port, keep the Manifest V3 service worker alive, and cancel cleanly.
- Author
- Extenify Team
- Published
- Reading time
- 12 min read

On this page
If you stream AI responses in a Chrome extension the same way you would on a web page, it works on your machine and then breaks for users. The content script cannot reach your API, sendMessage delivers only one reply, and the Manifest V3 service worker can be shut down halfway through a long answer. The result is the bug report every AI extension gets sooner or later: "the answer stops in the middle."
This guide shows a setup that holds up: a small backend that streams server-sent events (SSE), a service worker that relays tokens over a long-lived port, a content script that renders them safely, and the timing rules that keep the worker alive. The examples use the Claude Messages API and note what changes for the OpenAI Responses API.
Key takeaways
- Content scripts make requests as the page they run in, so they cannot call your API directly. Send the prompt to the service worker or an extension page and stream from there.
- Use
chrome.runtime.connect()ports, notsendMessage, for streaming. A port carries many messages, and since Chrome 114 sending messages on it keeps the service worker alive.- Chrome stops a service worker when "a fetch() response takes more than 30 seconds to arrive." Have your backend send headers and a first byte immediately, then a heartbeat while the model thinks.
- Abort the upstream request when the user closes the panel, so you stop streaming tokens nobody will read.
- Model output is untrusted text. Render it with
textContentor a sanitizer, and never execute it.
Why it is harder to stream AI responses in a Chrome extension
On a normal web page, streaming is one fetch() and a loop over the response body. In an extension, three platform rules get in the way.
Content scripts are not your extension's origin. Chrome's network requests documentation says that "content scripts initiate requests on behalf of the web origin that the content script has been injected into", and that "cross-origin requests are always treated as such in content scripts, even if the extension has host permissions." Your API is cross-origin for every site the user visits, so the call has to happen in the service worker or an extension page, which can reach any host listed in host_permissions.
One-off messages carry one reply. chrome.runtime.sendMessage gets a single response. The familiar error "The message port closed before a response was received" comes from this model: the channel closes once the listener is done. Streaming needs a channel that stays open, which is what long-lived connections are for.
The service worker is not always running. Chrome's service worker lifecycle page lists when it stops a worker:
- "After 30 seconds of inactivity. Receiving an event or calling an extension API resets this timer."
- "When a single request, such as an event or API call, takes longer than 5 minutes to process."
- "When a fetch() response takes more than 30 seconds to arrive."
It also lists what extends the lifetime, including, from Chrome 114, "sending a message with long-lived messaging keeps the service worker alive." Note the wording: opening a port alone "no longer resets the timers." Messages have to keep flowing.
EventSource does not help either. It only makes GET requests and cannot set an authorization header, so you cannot send a prompt with it. Use fetch() and read the body stream yourself.
Pick where the stream runs
| Where your UI lives | Where to run the fetch | Notes |
|---|---|---|
| Side panel or options page | In that page | Extension pages are not subject to the worker's idle rules. Simplest option. |
| Popup | In the popup, or the service worker | The popup closes when it loses focus, which kills a stream running in it |
| Content script UI on a web page | Service worker, relayed over a port | The pattern this guide builds |
If your feature lives in the side panel, you can skip the relay and run the same fetch() loop in the panel. Most AI extensions also want an inline UI next to a text field or selection, though, and that means a content script, a port and the service worker.
The offscreen documents API is sometimes suggested as a place to hold long streams. None of its documented reasons describes network requests, so treat that as a workaround rather than the intended use.
Step 1: Stream SSE from your backend
Your API key belongs on your server, not in the extension package, where anyone can unzip it. The server also gives you one place to authenticate users, rate limit, log usage and normalize each provider's event format into something small.
First, a parser for server-sent events that works in Node.js and in the extension. It reads a fetch() body and yields one event at a time:
// sse.js
export async function* readSse(body) {
const reader = body.pipeThrough(new TextDecoderStream()).getReader();
let buffer = '';
while (true) {
const { value, done } = await reader.read();
if (done) return;
buffer += value.replace(/\r\n/g, '\n');
let boundary;
while ((boundary = buffer.indexOf('\n\n')) !== -1) {
const block = buffer.slice(0, boundary);
buffer = buffer.slice(boundary + 2);
let event = 'message';
const data = [];
for (const line of block.split('\n')) {
if (line.startsWith('event:')) event = line.slice(6).trim();
else if (line.startsWith('data:')) data.push(line.slice(5).trimStart());
}
if (data.length === 0) continue; // comments and heartbeats
const text = data.join('\n');
let parsed = text;
try { parsed = JSON.parse(text); } catch { /* "[DONE]" and plain text */ }
yield { event, data: parsed };
}
}
}
Now the server. It sends headers right away, so the extension's fetch() resolves well inside the 30-second window, writes a heartbeat every 15 seconds while the model is quiet, and aborts the upstream call if the client disconnects.
// server.mjs (Node.js 22 or later, no dependencies besides sse.js)
import http from 'node:http';
import { readSse } from './sse.js';
const MODEL = process.env.AI_MODEL; // keep the model ID in config, not code
http.createServer(async (req, res) => {
if (req.method !== 'POST' || req.url !== '/v1/complete') {
res.writeHead(404).end();
return;
}
// Authenticate the user or install here before spending tokens.
let raw = '';
for await (const chunk of req) raw += chunk;
const { prompt } = JSON.parse(raw);
res.writeHead(200, {
'content-type': 'text/event-stream',
'cache-control': 'no-cache',
'x-accel-buffering': 'no', // tell nginx not to buffer the stream
});
res.write(': open\n\n'); // first byte now, not when the model starts
const upstream = new AbortController();
res.on('close', () => upstream.abort()); // client left: stop generating
const heartbeat = setInterval(() => res.write('event: ping\ndata: {}\n\n'), 15_000);
try {
const response = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
signal: upstream.signal,
headers: {
'content-type': 'application/json',
'x-api-key': process.env.ANTHROPIC_API_KEY,
'anthropic-version': '2023-06-01',
},
body: JSON.stringify({
model: MODEL,
max_tokens: 1024,
stream: true,
messages: [{ role: 'user', content: prompt }],
}),
});
if (!response.ok || !response.body) throw new Error(`Upstream HTTP ${response.status}`);
for await (const { data } of readSse(response.body)) {
if (data.type === 'content_block_delta' && data.delta?.type === 'text_delta') {
res.write(`data: ${JSON.stringify({ t: data.delta.text })}\n\n`);
} else if (data.type === 'error') {
throw new Error(data.error?.type ?? 'upstream_error');
}
}
res.write('data: [DONE]\n\n');
} catch (err) {
if (!upstream.signal.aborted) {
res.write(`event: error\ndata: ${JSON.stringify({ message: String(err.message) })}\n\n`);
}
} finally {
clearInterval(heartbeat);
res.end();
}
}).listen(8787);
A few details in that code matter more than they look:
- The event filter. Anthropic's streaming documentation describes a flow of
message_start, content blocks made ofcontent_block_start,content_block_deltaandcontent_block_stop, thenmessage_deltaandmessage_stop, withpingevents in between. It also warns that errors such asoverloaded_errorcan arrive inside the stream, and that "new event types may be added", so the code ignores anything it does not recognize. - OpenAI is the same shape. With the Responses API, set
stream: trueand forward thedeltafield ofresponse.output_text.deltaevents; the stream ends withresponse.completed. Because the server translates both into{ t }messages, the extension never needs to know which provider answered. - Proxy buffering. nginx buffers proxied responses by default, which turns a stream into one late block. The nginx proxy module documentation says buffering "can also be enabled or disabled by passing 'yes' or 'no' in the 'X-Accel-Buffering' response header field." Its default
proxy_read_timeoutis 60 seconds, which the 15-second heartbeat stays well under.
If your backend is due for a Node.js upgrade, our Node.js 26 LTS upgrade checklist covers the changes that affect servers like this one.
Step 2: Relay the stream through the service worker
Add your API host to the manifest so the service worker can call it:
{
"manifest_version": 3,
"background": { "service_worker": "background.js", "type": "module" },
"host_permissions": ["https://api.example.com/*"]
}
The service worker accepts a port, runs the fetch, and posts each chunk back. Every postMessage on the port counts as long-lived messaging, which is what keeps the worker running during the answer.
// background.js
import { readSse } from './sse.js';
const API_URL = 'https://api.example.com/v1/complete';
chrome.runtime.onConnect.addListener((port) => {
if (port.name !== 'ai-stream') return;
const controller = new AbortController();
port.onDisconnect.addListener(() => controller.abort());
port.onMessage.addListener(async ({ prompt }) => {
try {
const res = await fetch(API_URL, {
method: 'POST',
signal: controller.signal,
headers: { 'content-type': 'application/json' }, // add your auth header
body: JSON.stringify({ prompt }),
});
if (!res.ok || !res.body) throw new Error(`HTTP ${res.status}`);
for await (const { event, data } of readSse(res.body)) {
if (event === 'ping') port.postMessage({ type: 'ping' }); // keeps the worker alive
else if (event === 'error') throw new Error(data.message);
else if (data === '[DONE]') break;
else port.postMessage({ type: 'delta', text: data.t });
}
port.postMessage({ type: 'done' });
} catch (err) {
if (err.name !== 'AbortError') port.postMessage({ type: 'error', message: err.message });
}
});
});
Forwarding the heartbeat as a ping message is deliberate. Reasoning models can go quiet for a long time before the first token, and a stream with no messages on the port gives Chrome no reason to keep the worker running.
Step 3: Render the stream in the content script
The content script opens the port, sends the prompt, and appends text as it arrives:
// content.js
function streamCompletion(prompt, onText) {
return new Promise((resolve, reject) => {
const port = chrome.runtime.connect({ name: 'ai-stream' });
let finished = false;
port.onMessage.addListener((msg) => {
if (msg.type === 'delta') onText(msg.text);
else if (msg.type === 'done' || msg.type === 'error') {
finished = true;
port.disconnect();
msg.type === 'done' ? resolve() : reject(new Error(msg.message));
}
});
port.onDisconnect.addListener(() => {
if (!finished) reject(new Error('The answer was interrupted. Please try again.'));
});
port.postMessage({ prompt });
});
}
const output = document.createElement('div');
streamCompletion('Summarize this page', (text) => {
output.textContent += text; // never innerHTML with model output
}).catch((err) => {
output.dataset.state = 'error';
output.append(` ${err.message}`);
});
Two rules for the rendering side:
- Treat output as untrusted. A model can echo anything that was on the page, including markup that someone planted there. Use
textContent, or run Markdown through a sanitizer before inserting HTML. - Never execute it. Manifest V3 does not allow remotely hosted code, and text that arrives from your server and gets evaluated is exactly that. It is also the kind of finding that gets an item taken down; our guide on what to do when a Chrome extension is removed from the Chrome Web Store covers what happens next.
Step 4: Make it survive the real world
Once the happy path works, these are the cases that turn into support tickets:
- The user closes the tab or panel mid-answer. The port disconnects, the service worker aborts its fetch, the server sees the connection close and aborts the provider call. Test this path on purpose, because it is the one that saves money.
- The worker restarts anyway. Keep requests short enough that this is rare, and handle it when it happens: the content script's
onDisconnectfires without adone, so show a retry button with the partial text kept. If you need to resume, store the partial answer inchrome.storage.sessionas it arrives. - Very long answers. Google's lifecycle page also lists a five-minute limit for a single request such as an event or API call, and it does not spell out how that applies to a long stream. Bound
max_tokens, and test your longest realistic answer on the oldest Chrome version you support. - Provider errors and retired models. An
overloaded_errormid-stream or a model ID that no longer exists should produce a clear message, not a frozen cursor. Our guide to LLM model retirements in late 2026 shows how to keep model IDs out of the package so a shutdown does not need a store release.
Measure the streaming feature
Streaming changes how fast an answer feels, not how much it costs, so measure it from the user's side. Your server already sees every request, which means you can log most of this without adding tracking to the extension:
- Time to first token: from request received to first
tchunk written. - Completion rate: requests that reached
[DONE], versus client disconnects and upstream errors. - Errors by type: overloaded, rate limited, model not found, timeouts.
- Tokens per request: from the usage your provider reports at the end of the stream.
If you also want client-side events, such as how often people copy an answer, our guide to adding Google Analytics 4 to a Manifest V3 extension shows how to send them from the service worker without exposing a secret.
Then watch the store. An AI feature that stalls shows up in reviews and ratings within days, often before your own dashboards make the pattern obvious. Extenify's extension analytics dashboard records the public listing data for your extension in the Chrome Web Store, Edge Add-ons and Firefox Add-ons every day, with version markers on the growth chart and a daily change alert to email, Slack, Telegram, Discord or Teams when users or ratings move. Pair it with our checklist of what to watch in the week after an extension release. Tracking one extension is free; the pricing page has the rest.
Frequently asked questions
Why does my Chrome extension stop streaming in the middle of an answer?
Usually the service worker was stopped. Chrome stops it after 30 seconds without events or extension API calls, or when a fetch() response takes more than 30 seconds to arrive. Send headers and a first byte from your server immediately, forward a heartbeat while the model is quiet, and relay chunks over a port so messages keep flowing.
Can I use EventSource in a Chrome extension?
You can, but it rarely fits AI APIs. EventSource only makes GET requests and cannot set custom headers, so you cannot send a prompt body or an authorization header. Use fetch() with a POST and read the response body as a stream.
Can a content script call the AI API directly?
Not reliably. Content scripts make requests on behalf of the page they are injected into, and cross-origin requests are treated as cross-origin even when the extension has host permissions. Send the prompt to the service worker or an extension page and make the request there.
Should the extension call OpenAI or Claude directly with an API key?
No. Anything in the extension package can be read by anyone who downloads it, including API keys. Put the key on your own server, authenticate your users there, and stream from the server to the extension.
Why use a port instead of chrome.runtime.sendMessage?
sendMessage delivers one response per message, which cannot carry a stream of tokens. A port from chrome.runtime.connect() stays open for many messages in both directions, and since Chrome 114 sending messages on it keeps the service worker alive.
Does closing the panel stop the AI request?
Only if you wire it up. Abort the extension's fetch() when the port disconnects, and abort the provider call on your server when the client connection closes. Check your provider's documentation for how it bills a response that was cut off part way.
Summary
To stream AI responses in a Chrome extension, keep the network call out of the content script, relay the stream over a long-lived port, and treat the Manifest V3 service worker's timing rules as part of your design: headers within 30 seconds, messages flowing during long pauses, and a clean abort path from the tab to the provider. Render the output as text, never as code, and measure time to first token and completion rate on your server. Then watch your listing for the first reviews after release, because users notice a stalled answer before any dashboard does.
Image credits
- Cover photo: Stone water channel carrying a stream through a forest by Damián Ponte Mosqueira on the WordPress Photo Directory, CC0.


