Build your agent
Auto-index visited pages
Auto-index is a per-agent toggle that grows your knowledge base automatically as visitors browse your site. When a visitor lands on a page the agent has never seen, that page goes into the crawl queue — silently, in the background.
How to enable
Open Knowledge › Sources of the AI employee (/app/agents/{id}/sources)
and switch on Learn the pages visitors open (it stores
auto_index_visited_pages). The change takes effect
immediately — there's no separate publish step for this toggle.
What it does
On every /v1/widget/init call, after the agent passes the
origin and quota checks, AutoIndexPageVisit::attempt() runs
a chain of seven guards. All seven must pass for a crawl to be queued:
- The agent has
auto_index_visited_pages = true. - The page URL is a valid http/https URL.
- The host is not private — RFC1918 (
10.x,172.16-31.x,192.168.x), loopback (localhost,127.x,::1), link-local (169.254.x,fe80:),0.x, IPv6 ULA (fc00:), and.local/.internaldomains are blocked. - The path doesn't look private —
/admin,/login,/checkout,/profile,/account,/settings,/cart,/apiare skipped. - The URL hasn't already been indexed for this agent — by any
source, and with or without a trailing slash, so
/waterhardheid/meppelcounts as indexed when a sitemap stored/waterhardheid/meppel/. - We're under the rate limit — 30 crawls per agent per hour, counted in the cache.
- The visitor's actual
Originheader matches the agent'sallowed_origins(or matches the page URL's origin when allowed_origins is*).
If everything passes, we lazy-create the agent's one
type=auto source (if it doesn't exist yet) and dispatch a
CrawlPageJob for the page on the crawl queue. The
visitor's request returns immediately — auto-index never blocks the hot
path.
What it skips
The path blocklist exists because authenticated pages are noisy and
risky to index — a logged-in /profile or
/account/orders page leaks the visitor's data into your
knowledge base. The full list lives in
AutoIndexPageVisit::SKIP_PATH_PATTERNS (case-insensitive,
matches the path segment with optional trailing slash):
/account,/my-account,/profile,/settings/admin/login,/signin,/signup,/register,/logout,/auth,/password/checkout,/cart,/order,/orders
/portal, /customer), the default blocklist
won't catch them. Disable auto-index, or pre-list the exact public
URLs you want crawled and skip the toggle entirely.
Rate limiting
The 30-crawls-per-agent-per-hour cap is a counter keyed in the cache as
auto-index:agent:{id}:hour:{YmdH}. If a popular page on your
site is getting hammered the limit will quickly throttle, but normal
traffic patterns rarely hit it.
What gets indexed
Every agent has at most one type=auto source, shown in the
Sources list as Auto-indexed from visitors.
Each page it picks up becomes a document under that one source, so
you can preview or reindex them together; while it has collected
nothing yet it shows as Listening. A URL that already has a
document for this agent (guard 5 above) is not crawled again.
Before card #681, guard 5 compared the exact URL. The visit arrives in its canonical form (no trailing slash), while a sitemap or crawl source keeps the URL as fetched — so on a site whose URLs end in a slash, every page a visitor opened was crawled a second time. Those second copies may still be in your Auto-indexed from visitors source. They no longer affect answers: when two copies of one page match a question, only one of them is used (see Knowledge). Deleting the auto source removes them.
Disabling and pruning
Turn the toggle off and no new pages will be queued, but the pages already collected stay. To remove them, delete the Auto-indexed from visitors source; its documents and chunks go with it.