Skip to content
GitHub

Testing and improving

Try an agent in the dashboard before it reaches a channel, and fix what it gets wrong.


On this page

Every agent has an Improve & Try tab. It is a private conversation with the agent using its real instructions, knowledge, and tools, so what you see there is what a user in Slack would get.

Trying the agent#

Ask the questions you expect real users to ask, including the ones you are not sure it can answer. Past sessions are listed beside the chat, so you can compare an answer before and after a change instead of relying on memory.

Tool calls run for real. A message that asks the agent to create a ticket creates a ticket.

Reading a bad answer#

Most failures fall into three kinds, and each has a different fix.

What you seeUsual causeWhere to fix it
Confident but wrong factsThe source is not synced, or the answer is not in itKnowledge
Right facts, wrong shape or toneInstructions do not say what the answer should look likeInstructions
Did not act, or used the wrong toolThe tool is not connected, or its description is ambiguousTools

If the agent answers correctly here but not in a channel, the difference is usually channel configuration rather than the agent: check what triggers it on that channel. On Zendesk, also check whether the connection is restricted to replies the knowledge base can ground (Reply only when knowledge base has an answer).

After it is live#

Once real conversations are running, the Activity and Analytics tabs show what people actually ask and how they react. Activity holds the individual conversations and the trace behind each one; Analytics aggregates them. Agents on the Claude Agent SDK runtime have no Analytics tab. Emoji reactions on the agent's replies are counted as feedback on the Analytics tab, which is the cheapest signal you will get about answer quality. On Slack you can also turn on response feedback, which adds thumbs-up and thumbs-down buttons under each reply.

Analytics

The Classic Editor#

Anthropic and OpenAI Responses agents have a Classic Editor link at the bottom of the agent's sidebar. It opens the older editor, which some operations still need. For example, an OpenAI Responses agent whose uploaded files live in a vector store on your own OpenAI account switches back to Runbear's key from there, because the classic editor re-provisions the store. Claude Agent SDK and Classic agents don't have the link. See Runtimes.