B Blengi docs

Architecture

Crawler chain & vision OCR

Every URL-based source (url, sitemap, feed, auto) flows through the same multi-tier crawler chain. Each tier tries a different way to get readable text out of the page. The first tier whose extracted text clears the 200-character threshold wins; failures cascade to the next tier.

The chain (highest priority → last resort)

#TierWhat it doesWhen it wins
1 plain_http Plain Guzzle GET with a Mozilla user-agent. Runs through UrlSafetyGuard so private IPs and cloud metadata services are blocked. A response that is a PDF (declared application/pdf, or starting with the %PDF- bytes), including a page that redirects to one, is turned into HTML from the PDF's text with the same parser file uploads use, up to 20 MB (#642). The browser tiers render a PDF to an empty body, so this tier is the only one that can read it. When the site refuses the browser identity with 403 or 429 (a bot shield that blocks browsers coming from servers), the tier asks once more under its own name, {AppName}Bot/1.0 (+APP_URL), e.g. BlengiBot/1.0 (+https://blengi.com) (#644). robots.txt is read the way RFC 9309 describes it (#722): a group that names this crawler applies instead of the * group, Allow lines count, and so do * and $ in paths; the longest matching rule wins and Allow wins a tie. So a site closed to everyone can let only us in with User-agent: BlengiBot + Allow: /. The SSRF guard runs on every hop. A customer never has to allow our server's IP. Server-rendered pages (Blade apps, classic blogs, marketing sites with SSR or static export).
2 cloudflare_browser_markdown POST /browser-rendering/markdown. Cloudflare spins up headless Chrome, waits for JS to settle, and emits clean markdown. We wrap it in a minimal HTML envelope so the extractor and chunker stay unchanged. JS-heavy SPAs whose final DOM is text-shaped (most React / Vue / Svelte sites).
3 cloudflare_browser POST /browser-rendering/content. Same headless Chrome, but returns raw rendered HTML. Our ReadabilityExtractor takes over. Atypical layouts where the markdown extractor's heuristics drop content but the regex extractor catches it (tables that don't map cleanly to GFM, deep nested article structures).
4 cloudflare_vision Two-stage:
  1. Full-page screenshot via /browser-rendering/screenshot (PNG).
  2. Workers AI multimodal call to @cf/meta/llama-3.2-11b-vision-instruct with an OCR-only prompt. Returns extracted text.
Text is wrapped in <p> elements so the chunker treats each paragraph as its own unit.
Canvas-rendered slide decks, all-image landing pages, embedded PDF viewers, or any site whose meaningful content lives only in pixels.

How a tier "wins"

The chain measures extracted text length, not raw HTML length. Each tier's output is passed through ReadabilityExtractor (with the same chrome-stripping fallback used downstream). If the extracted text is at least 200 characters, that tier wins and its HTML is cached for 5 minutes against the source URL. If not, the chain falls forward.

This used to be a raw-length check, which let JS-only Inertia shells "win" with 50KB of empty <div> markup and then silently fail extraction downstream. The extraction-aware check fires the vision tier exactly when the extractor would have failed anyway — no wasted retries, no silent zeros.

Cost & safeguards

  • CLOUDFLARE_BROWSER_DAILY_LIMIT — shared counter for /content, /markdown, and /screenshot invocations. Default 0 (unlimited). Operators on a fixed Cloudflare bundle should set this to a per-day cap that matches their plan.
  • CLOUDFLARE_VISION_DAILY_LIMIT — separate counter for Workers AI vision calls (these consume Neurons, billed separately from Browser Rendering). Default 0.
  • CLOUDFLARE_VISION_MODEL — defaults to @cf/meta/llama-3.2-11b-vision-instruct. Override only if Cloudflare retires the model or you've negotiated a different one.
  • Both daily-cap guards throw a recoverable exception when the cap is hit, so the chain falls forward to the next tier instead of failing the whole crawl. Vision being capped just means the page that needed OCR doesn't get indexed — every other URL still flows.

What happens when every tier fails

The chain throws a RuntimeException whose message lists each tier and why it failed (HTTP code, exception message, or "extraction produced N chars under threshold"). CrawlPageJob handles it without throwing, because an unreachable customer site is not an error in this system and must not write ERROR lines that page the operator (#642):

  • Hard failures (DNS, connection refused, TLS, any 4xx except 429) fail at once; another try would get the same answer.
  • Everything else (429, 5xx, timeouts, empty renders) is retried after 30, 90 and 180 seconds (at least 60 for a 429), then failed. On the sync queue, which runs a job once, the first failure is final.
  • Failing records the job in failed_jobs and runs CrawlPageJob::failed(), which stamps the reason on the source's error column when the source has no documents (shown in the admin's Sources dialog through SourceErrorPresenter) and logs a crawl.page_unreachable warning.

A failed crawl keeps the page's last indexed copy

A rate limit, an outage or a firewall says nothing about the page itself, so the document indexed on the last successful crawl stays and keeps answering; the next refresh tries again. Only two outcomes remove it:

  • The page is gone: HTTP 404 or 410, or a page that answers 200 with "page not found" wording.
  • robots.txt now forbids it.

Gate pages (login wall, paywall, JavaScript-only shell, cookie wall, bot challenge) keep the last copy too: their wording overlaps with firewall block pages ("Access denied"). The bot-challenge check knows Cloudflare's 2026 interstitial ("Performing security verification", title "Just a moment..."); before, that page cleared the 200-character threshold and was indexed in place of the real page. Before #642 every failure deleted the page's document, so one night of a customer's host answering 429 to our server could have emptied their whole website source.

The www → apex certificate fallback

A very common site misconfiguration is a TLS certificate that covers the apex domain (example.com) but not its www host. Fetching https://www.example.com/… then fails every tier with cURL error 60 (SSL: no alternative certificate subject name matches target host name 'www.example.com') — and because that error is classed as a permanent failure, the page would never re-index. Left unhandled, the assistant answers from stale or missing content (a live case: a pricing page's "3,000 conversations / month" going unindexed).

So before giving up, the chain retries the apex host once: it strips the leading www. (preserving scheme, port, path and query) and re-runs the whole chain. The retry is gated strictly on the cert-mismatch signature — a genuinely broken www-only host (500, DNS failure, 404) is not handed a pointless second fetch — and it is self-terminating, since the apex URL has no www. prefix to fall back from. The recovery is logged as crawler.chain.www_apex_fallback.

Browserless deprecation

The BrowserlessClient tier was removed from the default chain in 2026-06 once the vision OCR tier landed. CF /markdown + /content + vision now cover the same ground free (no third-party billing line). The BROWSERLESS_URL / BROWSERLESS_TOKEN env vars stay readable for backwards compatibility — they no longer wire anything by default. A custom service provider can still bind BrowserlessClient manually if a self-hosted Browserless is preferred.