Skip to content
GitHub

Chat

Send messages to an agent through threads and runs or through chat completions, and read a conversation's history back.


On this page

A conversation with an agent lives in a thread. Each request you send runs one turn: the agent reads the thread, answers the newest message, and records its reply on the thread. There are two ways to run turns, and both share the same threads, limits and errors.

Two ways to chat#

Threads and runsChat completions
RequestsPOST /v1/threads, then POST /v1/threads/{threadId}/runs or /runs/streamPOST /v1/chat/completions
ThreadYou create it first and know its idCreated for you when you omit thread_id
StreamingA separate route, /runs/streamstream: true on the same route
Rolesuser, assistant, systemuser, assistant
Best forProducts that keep conversations and show them againOne-off questions and simple integrations

Both need an API key with the chat capability. See Authentication and API keys.

Threads and runs#

POST /v1/threads with { "assistant_id": "<agent id>" } answers 200 with { "thread": { "id": "<thread id>" } }.

POST /v1/threads/{threadId}/runs runs a turn and answers with the reply:

{
  "assistant_id": "<agent id>",
  "messages": [{ "role": "user", "content": "Summarize yesterday's escalations." }]
}
{ "message": { "id": "<message id>", "content": "Here are the three escalations..." } }

Send assistant_id on every run, and send the agent that owns the thread. POST /v1/threads/{threadId}/runs/stream takes the same body and answers with a stream.

Chat completions#

POST /v1/chat/completions takes assistant_id, messages, and optionally thread_id, stream and config. Without thread_id it creates a new thread for the request.

  • Without stream, it answers { "message": { "id", "content" }, "threadId" }.
  • With stream: true, it answers with a stream.

Either way, the response carries the thread's id in the x-runbear-thread-id header. Pass it back as thread_id to continue the conversation. The header is exposed to browsers through CORS.

Messages#

Each element of messages has role, content, an optional id and optional attachments. At least one message is required. Until a pending API fix ships, an empty messages array on POST /v1/chat/completions fails with 500 rather than 400, so check for it before you send.

  • The last message is the turn. It is what the agent answers, so send the new user message last.
  • Every message is recorded on the thread, and the agent reads the thread's recorded history. You can send only the new message each turn, or resend the whole conversation: a message that is already recorded is not added twice.
  • id is the deduplication key. Two messages with different ids are two messages; a repeated id is one. Without an id, Runbear derives one from the message itself, so two identical messages without ids collapse into one. Give each message its own id when a user can legitimately send the same text twice. On runs, an id is 1 to 200 characters.
  • The rbc: prefix is reserved for the Web SDK's response component clicks. On runs, an id that begins with rbc: must have the form rbc:<uuid>:<uuid>, or the request is refused with 400.

User context#

config.userContext is an object of string, number or boolean values. Runbear turns each value into a string and passes the object to the agent as metadata for the turn, where it can scope tool calls and knowledge retrieval, for example to the signed-in customer's account id.

{
  "assistant_id": "<agent id>",
  "messages": [{ "role": "user", "content": "Where is my order?" }],
  "config": { "userContext": { "customerId": "cus_1842", "plan": "pro" } }
}

Reading a conversation#

GET /v1/threads/{threadId}/messages returns the thread and its messages:

{
  "truncated": false,
  "thread": { "id": "<thread id>", "title": "Escalations", "assistantId": "<agent id>" },
  "messages": [
    { "id": "...", "role": "user", "content": "Summarize yesterday's escalations.", "traceId": null },
    { "id": "...", "role": "assistant", "content": "Here are the three escalations...", "traceId": "..." }
  ]
}
  • A long thread returns only its newest 200 messages, with truncated: true. There is no pagination for older messages.
  • traceId links a reply to its trace.
  • An assistant message with interrupted: true holds only the text produced before its turn ended without delivering a reply.
  • components lists the response components a reply carried, when it carried any.

A thread that doesn't exist and a thread in another organization answer the same 404, so the response never reveals whether an id exists.

Listing threads#

GET /v1/threads?assistant_id=<agent id> lists the agent's threads, newest first, as { "threads": [{ "id", "createdAt", "title" }], "nextCursor" }.

  • It returns only threads created through the API for that agent, not conversations from Slack, Teams or other channels.
  • limit is 1 to 100, and defaults to 100.
  • Pass nextCursor back as cursor for the next page; it is null on the last page. Treat cursors as opaque. A cursor longer than 1024 characters, or one the API did not issue, is refused with 400.

How long a turn can run#

Each turn is bounded by the agent's timeoutMinutes setting: 6 minutes by default and at most 20. Claude Agent SDK agents are capped at 15 minutes. A turn that runs out of time ends with a timeout stream error.

Agents the API can't run#

  • An agent of type openai-assistant is refused with 422 (Agents of type "openai-assistant" are no longer usable.), and so is an agent whose type the API doesn't recognize. The refusal comes before a stream starts, so streaming routes answer 422 too.
  • A turn that fails permanently after it has started answers 422 with { "error": "unprocessable_entity", "message" } on non-streaming routes. On streaming routes the response has already begun with 200, so the failure arrives as a fatal error event instead.

Suggested follow-ups#

POST /v1/chat/suggestions takes assistant_id and messages (without attachments) and returns { "suggestions": ["...", "..."] }: short prompts a user might send next. It answers null when the agent isn't found. Each call counts against the chat rate limit, and it isn't available with a Web SDK session pass.