B Blengi docs

Build your agent

Auto-index visited pages

Auto-index is a per-agent toggle that grows your knowledge base automatically as visitors browse your site. When a visitor lands on a page the agent has never seen, that page goes into the crawl queue — silently, in the background.

How to enable

Open Knowledge › Sources of the AI employee (/app/agents/{id}/sources) and switch on Learn the pages visitors open (it stores auto_index_visited_pages). The change takes effect immediately — there's no separate publish step for this toggle.

What it does

On every /v1/widget/init call, after the agent passes the origin and quota checks, AutoIndexPageVisit::attempt() runs a chain of seven guards. All seven must pass for a crawl to be queued:

  1. The agent has auto_index_visited_pages = true.
  2. The page URL is a valid http/https URL.
  3. The host is not private — RFC1918 (10.x, 172.16-31.x, 192.168.x), loopback (localhost, 127.x, ::1), link-local (169.254.x, fe80:), 0.x, IPv6 ULA (fc00:), and .local / .internal domains are blocked.
  4. The path doesn't look private — /admin, /login, /checkout, /profile, /account, /settings, /cart, /api are skipped.
  5. The URL hasn't already been indexed for this agent — by any source, and with or without a trailing slash, so /waterhardheid/meppel counts as indexed when a sitemap stored /waterhardheid/meppel/.
  6. We're under the rate limit — 30 crawls per agent per hour, counted in the cache.
  7. The visitor's actual Origin header matches the agent's allowed_origins (or matches the page URL's origin when allowed_origins is *).

If everything passes, we lazy-create the agent's one type=auto source (if it doesn't exist yet) and dispatch a CrawlPageJob for the page on the crawl queue. The visitor's request returns immediately — auto-index never blocks the hot path.

What it skips

The path blocklist exists because authenticated pages are noisy and risky to index — a logged-in /profile or /account/orders page leaks the visitor's data into your knowledge base. The full list lives in AutoIndexPageVisit::SKIP_PATH_PATTERNS (case-insensitive, matches the path segment with optional trailing slash):

  • /account, /my-account, /profile, /settings
  • /admin
  • /login, /signin, /signup, /register, /logout, /auth, /password
  • /checkout, /cart, /order, /orders
Don't index logged-in surfaces
If your app's auth-gated pages live under unusual paths (e.g. /portal, /customer), the default blocklist won't catch them. Disable auto-index, or pre-list the exact public URLs you want crawled and skip the toggle entirely.

Rate limiting

The 30-crawls-per-agent-per-hour cap is a counter keyed in the cache as auto-index:agent:{id}:hour:{YmdH}. If a popular page on your site is getting hammered the limit will quickly throttle, but normal traffic patterns rarely hit it.

What gets indexed

Every agent has at most one type=auto source, shown in the Sources list as Auto-indexed from visitors. Each page it picks up becomes a document under that one source, so you can preview or reindex them together; while it has collected nothing yet it shows as Listening. A URL that already has a document for this agent (guard 5 above) is not crawled again.

Before card #681, guard 5 compared the exact URL. The visit arrives in its canonical form (no trailing slash), while a sitemap or crawl source keeps the URL as fetched — so on a site whose URLs end in a slash, every page a visitor opened was crawled a second time. Those second copies may still be in your Auto-indexed from visitors source. They no longer affect answers: when two copies of one page match a question, only one of them is used (see Knowledge). Deleting the auto source removes them.

Disabling and pruning

Turn the toggle off and no new pages will be queued, but the pages already collected stay. To remove them, delete the Auto-indexed from visitors source; its documents and chunks go with it.