# Streaming

> Read an agent's reply as it is written, with the NDJSON event format, the full event list, error codes, and reference parsers in JavaScript and Python.

Source: https://docs.runbear.io/api/streaming

Last updated: 2026-09-30

A streaming request returns the agent's reply while it is being written, along
with the tools it calls and any interactive components it produces. Use it to
show text as it arrives instead of waiting for the whole turn.

Two routes stream:

- `POST /v1/threads/{threadId}/runs/stream`
- `POST /v1/chat/completions` with `"stream": true`

Both take the same request body as their non-streaming form. See [Chat](/api/chat.md).

## The format is NDJSON, not SSE

The response is sent with `Content-Type: text/event-stream`, but the body is
**newline-delimited JSON (NDJSON)**: one complete JSON object per line, each
ending in `\n`. There is no `data:` prefix and no blank line between events, so
Server-Sent Events parsers don't work. `EventSource` can't be used either,
because both routes are `POST`.

```text
{"event":"thread.message.delta","data":{"content":"Here are"}}
{"event":"thread.message.delta","data":{"content":"Here are the three escalations"}}
{"event":"thread.message.completed","data":{"id":"<message id>","content":"Here are the three escalations..."}}
{"event":"done"}
```

To read it, split the body on `\n` and parse each non-empty line. A line can be
split across two network reads, so keep any incomplete remainder and prepend it
to the next read. See the [reference parsers](#reference-parsers) below.

The response also carries `cache-control: no-cache, no-store, must-revalidate,
no-transform` and `x-accel-buffering: no`. Any proxy between your client and
Runbear must pass the stream through without buffering it, or events arrive all
at once at the end.

## Events

Every line has an `event` and, except for `done`, a `data` object.

| `event`                         | `data`                                                      | Meaning                                                                                                                                                           |
| ------------------------------- | ----------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `thread.message.delta`          | `{ content }`                                               | The reply text so far                                                                                                                                             |
| `thread.message.thinking_delta` | `{ content }`                                               | The model's reasoning so far, on agents that expose it                                                                                                            |
| `thread.tool_call.progress`     | `{ toolUseId, toolName, toolDisplayName, status, message }` | A tool call's state: `preparing`, `running`, `completed` or `failed`. Not sent when the agent's [tool activity](/agents/settings.md#tool-activity) setting is off |
| `thread.message.component`      | A component: `{ componentId, name, props, fallbackText }`   | An interactive [response component](/api/response-components.md) for the current reply                                                                            |
| `thread.message.completed`      | `{ id?, content }`                                          | A finished message and its final text                                                                                                                             |
| `error`                         | `{ error, code, fatal? }`                                   | A problem during the turn. See [Errors](#errors)                                                                                                                  |
| `done`                          | none                                                        | The stream is over                                                                                                                                                |

Treat any event you don't recognize as ignorable: new events may be added.

### Deltas replace, they don't append

`thread.message.delta` and `thread.message.thinking_delta` carry the **whole
text so far**, not the next piece. Replace what you are showing with each one
instead of appending to it.

The text is not guaranteed to only grow. A delta can briefly include tool-status
text that a later delta drops, so always render the latest value as-is.
`thread.message.completed` holds the final text.

### Components

A `thread.message.component` event belongs to the reply you are rendering, not
to the last completed message: depending on the tool that produced it, it can
arrive before or after `thread.message.completed`. Show `fallbackText` when you
don't render a component's `name`.

## How a stream ends

A stream normally ends with `{"event":"done"}`. A stream that ends without
`done` was cut off, for example by a network failure or a proxy timeout. The
turn may still have finished on Runbear's side, so read the thread's messages to
see what was recorded before you retry.

## Errors

A failure that happens before the stream starts, such as an invalid body, a
missing thread, or an agent the API can't run, is an ordinary HTTP error. See
[Errors and limits](/api/errors-and-limits.md). Once the stream has started, the status
is already `200`, so problems arrive as `error` events:

```json
{"event":"error","data":{"error":"...","code":"context_length_exceeded","fatal":true}}
```

- `error` is text written for a person. Don't parse it.
- `code` is the machine-readable cause. Branch on it.
- `fatal: true` marks a failure that ended the turn. `authorization_required`,
  `monthly_credit_limit_exceeded` and `organization_suspended` arrive without
  `fatal`. Every stream still ends with `done` unless it is cut off.

| `code`                          | Meaning                                                                  | What to do                                                  |
| ------------------------------- | ------------------------------------------------------------------------ | ----------------------------------------------------------- |
| `context_length_exceeded`       | The thread no longer fits the model                                      | Start a new thread; retrying cannot help                    |
| `provider_overloaded`           | The model provider is temporarily over capacity                          | Retry                                                       |
| `provider_authentication_error` | The provider rejected the credentials, such as your own provider key     | Fix the agent's provider key                                |
| `invalid_tool_schema`           | A tool attached to the agent has an invalid definition                   | Fix or remove the tool; a new thread cannot help            |
| `rate_limit`                    | The model provider rate-limited the request                              | Retry with backoff                                          |
| `timeout`                       | The turn exceeded the agent's timeout                                    | Retry, or raise the agent's timeout                         |
| `empty_response`                | The turn ended without a reply                                           | Retry                                                       |
| `authorization_required`        | A tool needs the user to connect an account                              | See below                                                   |
| `monthly_credit_limit_exceeded` | The organization used its monthly credits                                | Add credits or wait for the next period                     |
| `organization_suspended`        | The organization is suspended                                            | Contact Runbear support                                     |
| `assistant_not_found`           | The agent doesn't exist or can't be reached                              | Check the agent id                                          |
| `provider_error`                | The model provider returned an internal server error                     | Retry                                                       |
| `provider_credit_exhausted`     | Your own provider key has no credit left                                 | Top up the provider account, then retry                     |
| `empty_message`                 | The message had no usable content                                        | Send a message with content                                 |
| `max_output_tokens_reached`     | The model hit the agent's output-token limit before finishing            | Raise the limit in the agent's settings                     |
| `unsupported_model_parameter`   | The agent's model rejects one of its sampling settings                   | Change the agent's model settings; a new thread cannot help |
| `file_unavailable`              | The provider couldn't read a file or image in the thread                 | Upload it again in a new thread                             |
| `conversation_unrecoverable`    | The thread history can no longer be sent to the provider                 | Start a new thread                                          |
| `provider_request_rejected`     | The provider rejected the request for a reason Runbear couldn't classify | Retry later                                                 |
| `internal_error`                | Anything else                                                            | Retry once, then report it                                  |

Not every runtime reports every code yet. The more specific codes, such as
`provider_error` or `file_unavailable`, come from agents on the
[Anthropic runtime](/agents/runtimes.md). Other runtimes report the same failure
under a broader code, often `internal_error`.

The list only grows. Treat a `code` you don't recognize as `internal_error`.

### Connecting an account

`authorization_required` arrives without `fatal`. Its `error` text contains one
Markdown link per app to connect, in the form `[Connect <App>](<url>)`. Show the
text, or extract the links, so the user can connect the account. After
connecting, the user sends a new message; the turn that asked is not resumed.

## Reference parsers

Both parsers keep the incomplete end of each read in a buffer, parse only whole
lines, split only on `\n`, and parse the final line even when the stream ends
without a trailing newline.

### JavaScript

Works in browsers and in Node.js 18 and later.

```js
async function* readRunbearStream(response) {
  if (!response.ok) {
    throw new Error(`Runbear API error ${response.status}: ${await response.text()}`)
  }
  const reader = response.body.getReader()
  const decoder = new TextDecoder()
  let buffer = ""
  while (true) {
    const { value, done } = await reader.read()
    if (done) break
    buffer += decoder.decode(value, { stream: true })
    let newline = buffer.indexOf("\n")
    while (newline !== -1) {
      const line = buffer.slice(0, newline).trim()
      buffer = buffer.slice(newline + 1)
      if (line !== "") yield JSON.parse(line)
      newline = buffer.indexOf("\n")
    }
  }
  buffer += decoder.decode()
  const last = buffer.trim()
  if (last !== "") yield JSON.parse(last)
}

const response = await fetch(
  `https://api.runbear.io/v1/threads/${threadId}/runs/stream`,
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      assistant_id: agentId,
      messages: [{ role: "user", content: "What changed this week?" }],
    }),
  },
)

let reply = ""
for await (const event of readRunbearStream(response)) {
  if (event.event === "thread.message.delta") {
    reply = event.data.content // replace, don't append
  } else if (event.event === "thread.message.completed") {
    reply = event.data.content
  } else if (event.event === "error") {
    console.error(event.data.code, event.data.error)
  } else if (event.event === "done") {
    break
  }
}
```

### Python

Uses [httpx](https://www.python-httpx.org/). The parser splits on `\n` itself
rather than calling `iter_lines()`, because Python's line splitting also breaks
on characters such as `U+2028` that can appear inside a reply.

```python
import json
import httpx

def read_runbear_stream(chunks):
    buffer = b""
    for chunk in chunks:
        buffer += chunk
        *lines, buffer = buffer.split(b"\n")
        for line in lines:
            if line.strip():
                yield json.loads(line)
    if buffer.strip():
        yield json.loads(buffer)

body = {
    "assistant_id": agent_id,
    "messages": [{"role": "user", "content": "What changed this week?"}],
}

with httpx.stream(
    "POST",
    f"https://api.runbear.io/v1/threads/{thread_id}/runs/stream",
    headers={"Authorization": f"Bearer {api_key}"},
    json=body,
    timeout=httpx.Timeout(10.0, read=None),
) as response:
    response.raise_for_status()
    reply = ""
    for event in read_runbear_stream(response.iter_bytes()):
        if event["event"] in ("thread.message.delta", "thread.message.completed"):
            reply = event["data"]["content"]  # replace, don't append
        elif event["event"] == "error":
            print(event["data"]["code"], event["data"]["error"])
        elif event["event"] == "done":
            break
```

`read=None` turns off httpx's read timeout, because a long tool call can keep
the stream quiet for a while. With `requests`, send the request with
`stream=True` and pass `response.iter_content(chunk_size=None)` to the same
parser.

## Related

- [Chat](/api/chat.md) — the request body and the non-streaming routes
- [Response Components API](/api/response-components.md)
- [Errors and limits](/api/errors-and-limits.md)
