Streaming
Read an agent's reply as it is written, with the NDJSON event format, the full event list, error codes, and reference parsers in JavaScript and Python.
On this page
A streaming request returns the agent's reply while it is being written, along with the tools it calls and any interactive components it produces. Use it to show text as it arrives instead of waiting for the whole turn.
Two routes stream:
POST /v1/threads/{threadId}/runs/streamPOST /v1/chat/completionswith"stream": true
Both take the same request body as their non-streaming form. See Chat.
The format is NDJSON, not SSE#
The response is sent with Content-Type: text/event-stream, but the body is
newline-delimited JSON (NDJSON): one complete JSON object per line, each
ending in \n. There is no data: prefix and no blank line between events, so
Server-Sent Events parsers don't work. EventSource can't be used either,
because both routes are POST.
{"event":"thread.message.delta","data":{"content":"Here are"}}
{"event":"thread.message.delta","data":{"content":"Here are the three escalations"}}
{"event":"thread.message.completed","data":{"id":"<message id>","content":"Here are the three escalations..."}}
{"event":"done"}To read it, split the body on \n and parse each non-empty line. A line can be
split across two network reads, so keep any incomplete remainder and prepend it
to the next read. See the reference parsers below.
The response also carries cache-control: no-cache, no-store, must-revalidate, no-transform and x-accel-buffering: no. Any proxy between your client and
Runbear must pass the stream through without buffering it, or events arrive all
at once at the end.
Events#
Every line has an event and, except for done, a data object.
event | data | Meaning |
|---|---|---|
thread.message.delta | { content } | The reply text so far |
thread.message.thinking_delta | { content } | The model's reasoning so far, on agents that expose it |
thread.tool_call.progress | { toolUseId, toolName, toolDisplayName, status, message } | A tool call's state: preparing, running, completed or failed. Not sent when the agent's tool activity setting is off |
thread.message.component | A component: { componentId, name, props, fallbackText } | An interactive response component for the current reply |
thread.message.completed | { id?, content } | A finished message and its final text |
error | { error, code, fatal? } | A problem during the turn. See Errors |
done | none | The stream is over |
Treat any event you don't recognize as ignorable: new events may be added.
Deltas replace, they don't append#
thread.message.delta and thread.message.thinking_delta carry the whole
text so far, not the next piece. Replace what you are showing with each one
instead of appending to it.
The text is not guaranteed to only grow. A delta can briefly include tool-status
text that a later delta drops, so always render the latest value as-is.
thread.message.completed holds the final text.
Components#
A thread.message.component event belongs to the reply you are rendering, not
to the last completed message: depending on the tool that produced it, it can
arrive before or after thread.message.completed. Show fallbackText when you
don't render a component's name.
How a stream ends#
A stream normally ends with {"event":"done"}. A stream that ends without
done was cut off, for example by a network failure or a proxy timeout. The
turn may still have finished on Runbear's side, so read the thread's messages to
see what was recorded before you retry.
Errors#
A failure that happens before the stream starts, such as an invalid body, a
missing thread, or an agent the API can't run, is an ordinary HTTP error. See
Errors and limits. Once the stream has started, the status
is already 200, so problems arrive as error events:
{"event":"error","data":{"error":"...","code":"context_length_exceeded","fatal":true}}erroris text written for a person. Don't parse it.codeis the machine-readable cause. Branch on it.fatal: truemarks a failure that ended the turn.authorization_required,monthly_credit_limit_exceededandorganization_suspendedarrive withoutfatal. Every stream still ends withdoneunless it is cut off.
code | Meaning | What to do |
|---|---|---|
context_length_exceeded | The thread no longer fits the model | Start a new thread; retrying cannot help |
provider_overloaded | The model provider is temporarily over capacity | Retry |
provider_authentication_error | The provider rejected the credentials, such as your own provider key | Fix the agent's provider key |
invalid_tool_schema | A tool attached to the agent has an invalid definition | Fix or remove the tool; a new thread cannot help |
rate_limit | The model provider rate-limited the request | Retry with backoff |
timeout | The turn exceeded the agent's timeout | Retry, or raise the agent's timeout |
empty_response | The turn ended without a reply | Retry |
authorization_required | A tool needs the user to connect an account | See below |
monthly_credit_limit_exceeded | The organization used its monthly credits | Add credits or wait for the next period |
organization_suspended | The organization is suspended | Contact Runbear support |
assistant_not_found | The agent doesn't exist or can't be reached | Check the agent id |
provider_error | The model provider returned an internal server error | Retry |
provider_credit_exhausted | Your own provider key has no credit left | Top up the provider account, then retry |
empty_message | The message had no usable content | Send a message with content |
max_output_tokens_reached | The model hit the agent's output-token limit before finishing | Raise the limit in the agent's settings |
unsupported_model_parameter | The agent's model rejects one of its sampling settings | Change the agent's model settings; a new thread cannot help |
file_unavailable | The provider couldn't read a file or image in the thread | Upload it again in a new thread |
conversation_unrecoverable | The thread history can no longer be sent to the provider | Start a new thread |
provider_request_rejected | The provider rejected the request for a reason Runbear couldn't classify | Retry later |
internal_error | Anything else | Retry once, then report it |
Not every runtime reports every code yet. The more specific codes, such as
provider_error or file_unavailable, come from agents on the
Anthropic runtime. Other runtimes report the same failure
under a broader code, often internal_error.
The list only grows. Treat a code you don't recognize as internal_error.
Connecting an account#
authorization_required arrives without fatal. Its error text contains one
Markdown link per app to connect, in the form [Connect <App>](<url>). Show the
text, or extract the links, so the user can connect the account. After
connecting, the user sends a new message; the turn that asked is not resumed.
Reference parsers#
Both parsers keep the incomplete end of each read in a buffer, parse only whole
lines, split only on \n, and parse the final line even when the stream ends
without a trailing newline.
JavaScript#
Works in browsers and in Node.js 18 and later.
async function* readRunbearStream(response) {
if (!response.ok) {
throw new Error(`Runbear API error ${response.status}: ${await response.text()}`)
}
const reader = response.body.getReader()
const decoder = new TextDecoder()
let buffer = ""
while (true) {
const { value, done } = await reader.read()
if (done) break
buffer += decoder.decode(value, { stream: true })
let newline = buffer.indexOf("\n")
while (newline !== -1) {
const line = buffer.slice(0, newline).trim()
buffer = buffer.slice(newline + 1)
if (line !== "") yield JSON.parse(line)
newline = buffer.indexOf("\n")
}
}
buffer += decoder.decode()
const last = buffer.trim()
if (last !== "") yield JSON.parse(last)
}
const response = await fetch(
`https://api.runbear.io/v1/threads/${threadId}/runs/stream`,
{
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
assistant_id: agentId,
messages: [{ role: "user", content: "What changed this week?" }],
}),
},
)
let reply = ""
for await (const event of readRunbearStream(response)) {
if (event.event === "thread.message.delta") {
reply = event.data.content // replace, don't append
} else if (event.event === "thread.message.completed") {
reply = event.data.content
} else if (event.event === "error") {
console.error(event.data.code, event.data.error)
} else if (event.event === "done") {
break
}
}Python#
Uses httpx. The parser splits on \n itself
rather than calling iter_lines(), because Python's line splitting also breaks
on characters such as U+2028 that can appear inside a reply.
import json
import httpx
def read_runbear_stream(chunks):
buffer = b""
for chunk in chunks:
buffer += chunk
*lines, buffer = buffer.split(b"\n")
for line in lines:
if line.strip():
yield json.loads(line)
if buffer.strip():
yield json.loads(buffer)
body = {
"assistant_id": agent_id,
"messages": [{"role": "user", "content": "What changed this week?"}],
}
with httpx.stream(
"POST",
f"https://api.runbear.io/v1/threads/{thread_id}/runs/stream",
headers={"Authorization": f"Bearer {api_key}"},
json=body,
timeout=httpx.Timeout(10.0, read=None),
) as response:
response.raise_for_status()
reply = ""
for event in read_runbear_stream(response.iter_bytes()):
if event["event"] in ("thread.message.delta", "thread.message.completed"):
reply = event["data"]["content"] # replace, don't append
elif event["event"] == "error":
print(event["data"]["code"], event["data"]["error"])
elif event["event"] == "done":
breakread=None turns off httpx's read timeout, because a long tool call can keep
the stream quiet for a while. With requests, send the request with
stream=True and pass response.iter_content(chunk_size=None) to the same
parser.
Related#
- Chat — the request body and the non-streaming routes
- Response Components API
- Errors and limits