Websites
Knowledge base

Websites

Crawl your site into the knowledge base — review pages before crawling, sync on a schedule, and manage every page.

Website sources let the agent answer from your site. Add a URL, pick which pages to include, and Asks crawls them into the knowledge base. Manage them at app.asks.app/knowledge-base/websites.

Website URLs must be https — plain http URLs (and URLs with embedded credentials) are rejected. The site must be publicly reachable; pages behind a login can't be crawled. The one exception is a password-protected Shopify storefront — see Password-protected Shopify stores.

Add a website

Add the website

Click Add website and enter the top-level URL (for example https://example.com). Asks discovers subpages from there — via sitemap.xml by default. The dialog splits its settings across a Settings and an Advanced tab:

SettingTabWhat it doesDefault
Use sitemap.xmlSettingsDiscover pages from the site's sitemap automaticallyOn
Auto-refreshSettingsRe-sync cadence: never, daily (Business and Premium), weekly, or monthlyWeekly
Crawl DepthAdvancedMaximum link depth from the root URL (1–10)3
Allowed DomainsAdvancedComma-separated domains to stay within; empty allows all subdomains of the rootEmpty
Storefront passwordAdvancedPassword of a password-protected Shopify store, so Asks can read the siteEmpty
Configure Website dialog on the Settings tab with Website URL field, Use sitemap.xml toggle, and Auto-refresh cadence select, with an Advanced tab next to the Settings tab
Website settings. The Advanced tab holds crawl depth, allowed domains, and the storefront password.
Review pages before crawling

Saving opens the Choose pages to include picker. Asks maps the site's links first — without crawling anything — so you decide exactly which pages go in before any crawling starts. The link map covers up to 5,000 URLs per site, and a re-map button next to the search box refreshes it on demand.

Choose pages to include dialog listing mapped URLs grouped by path with checkboxes, a search box, a re-map button, and a selected-count footer
The review picker. On a brand-new website everything selectable starts checked, up to the per-crawl cap; later visits are seeded with your saved selection.

On Shopify stores, product, collection, cart, and checkout pages appear greyed out with a note: they're answered live through your Shopify connection and never use knowledge base articles, so they can't be selected here. Policy pages are excluded the same way by default. Machine-facing files (sitemaps, AI-crawler manifests) are auto-excluded too.

If you change your mind while adding a new site, Discard website removes the source without crawling anything.

Save and sync

Save your selection and the first sync starts immediately. The page count grows live while it crawls; the row shows page count and last-synced time when it finishes.

Each crawled page counts as one Knowledge Base Article — that's the only plan limit that applies, and re-crawling pages you already have never uses new articles. Each crawl ingests at most 2,000 pages per website (the same on every plan; the picker enforces the same cap on selections), and a URL already indexed under another website in your workspace is skipped rather than counted twice.

Password-protected Shopify stores

Development and pre-launch Shopify stores sit behind a storefront password. Asks can still crawl them: enter the password on the config dialog's Advanced tab — it's in Shopify under Online Store → Preferences — and Asks reads past the password page.

If a crawl runs into a password page it can't get past, the website row shows Password-protected store (or Store password not accepted if the saved password stopped working) with an Enter password prompt. The review picker shows the same notice and lists the store's real pages once the password is saved.

Syncing and re-crawls

Press Sync on a website row to re-crawl it, or let auto-refresh handle it on schedule. Re-syncs are incremental: every page's content is hashed, and pages whose content hasn't changed are skipped — nothing is re-indexed. Only genuinely new pages use new articles; changed pages refresh in place. Every sync also refreshes the site's link map, so pages added to your site since the last sync become selectable in the picker.

Syncs remove pages as well as add them: a healthy full crawl deletes pages that have disappeared from the site, and a selected-pages sync deletes pages you deselected. Removed pages free their article slots immediately.

Your plan sets how often each website can re-sync — weekly on Starter, daily on Business and Premium. The same frequency applies to the auto-refresh schedule and the manual Sync button (which shows when the next sync is available). A newly added website is exempt for its first 24 hours, so you can crawl, review pages, and re-sync freely while setting it up. Workspace-wide, crawls are also capped at 30 syncs per hour; beyond that, syncs are refused with "Too many website syncs in a short window" until the window passes.

Sync history (in the row's ⋯ menu) lists the last 50 syncs with status, start time, duration, pages crawled, and any error message, plus the source's success rate. If a sync goes wrong, the row says so: "Last sync incomplete" when it stopped early (for example on the article limit), "Last sync failed", or "Last sync finished with warnings".

Managing pages

Click View pages on a website to open its page list. From there you can:

  • Search pages by URL or title.
  • Re-crawl or delete pages individually or in bulk (up to 50 per action).
  • Open a page to see exactly what was extracted — a read-only view of the synced content with its URL, extracted size, an Open original page link, a Re-crawl this page action, and Delete page. Page content is synced from the website, so it can't be edited.
Website page list showing crawled pages with checkboxes, search input, and bulk re-crawl and delete actions
Per-page management. Deleting pages frees their article slots immediately.

Deleting a page stops the agent answering from it, but a future full sync may re-add it if it's still on the website — deselect it in Choose pages to keep it out for good. Choose pages (in the website's ⋯ menu) opens the same picker as at setup, seeded with your current selection; clearing the selection returns the source to automatic crawling.

Deleting a website

Delete in the ⋯ menu removes the source and all its crawled pages, and frees the articles they used. This can't be undone; the site can always be re-added and re-crawled later.