Skip to content
GitHub

Website

Crawl a public website into an agent's knowledge.


On this page

Crawl and sync content from any public website so your agent can answer questions grounded in that site's content.

How it works#

You provide one or more URLs. For each one, Runbear crawls that page and every subpage whose URL begins with it, extracts the text content, and indexes it for retrieval. For example, https://example.com/docs also picks up https://example.com/docs/getting-started, but not https://example.com/pricing. The agent searches this content at query time to provide accurate answers.

No authentication is required — Website works with any publicly accessible URL.

Setup#

  1. Open your agent in the Runbear dashboard and go to the Knowledge tab.
  2. Select Website.
  3. Enter a Website URL and click Add URL. Repeat for each site or section you want to include.
  4. Review the pages to include, then click Confirm. Runbear begins crawling and indexing them.

What gets synced#

  • Text content from each URL you provide and the subpages under it
  • Runbear re-crawls daily to pick up content changes

Limits#

  • Up to 200 subpages per website. Contact support if you need to sync more.
  • A page that fails with a lasting error, such as a 404 or a site that blocks the crawler, is retried after 7 days rather than on every daily re-crawl.