Websites
Websites
Crawl your site into the knowledge base — review pages before crawling, sync on a schedule, and manage every page.
Website sources let the agent answer from your site. Add a URL, pick which pages to include, and Asks crawls them into the knowledge base. Manage them at app.asks.app/knowledge-base/websites.
https — plain http URLs (and URLs with embedded credentials) are rejected. The site must be publicly reachable; pages behind a login can't be crawled. The one exception is a password-protected Shopify storefront — see Password-protected Shopify stores.Add a website
Click Add website and enter the top-level URL (for example https://example.com). Asks discovers subpages from there — via sitemap.xml by default. The dialog splits its settings across a Settings and an Advanced tab:
| Setting | Tab | What it does | Default |
|---|---|---|---|
| Use sitemap.xml | Settings | Discover pages from the site's sitemap automatically | On |
| Auto-refresh | Settings | Re-sync cadence: never, daily (Business and Premium), weekly, or monthly | Weekly |
| Crawl Depth | Advanced | Maximum link depth from the root URL (1–10) | 3 |
| Allowed Domains | Advanced | Comma-separated domains to stay within; empty allows all subdomains of the root | Empty |
| Storefront password | Advanced | Password of a password-protected Shopify store, so Asks can read the site | Empty |

Saving opens the Choose pages to include picker. Asks maps the site's links first — without crawling anything — so you decide exactly which pages go in before any crawling starts. The link map covers up to 5,000 URLs per site, and a re-map button next to the search box refreshes it on demand.

On Shopify stores, product, collection, cart, and checkout pages appear greyed out with a note: they're answered live through your Shopify connection and never use knowledge base articles, so they can't be selected here. Policy pages are excluded the same way by default. Machine-facing files (sitemaps, AI-crawler manifests) are auto-excluded too.
If you change your mind while adding a new site, Discard website removes the source without crawling anything.
Save your selection and the first sync starts immediately. The page count grows live while it crawls; the row shows page count and last-synced time when it finishes.
Password-protected Shopify stores
Development and pre-launch Shopify stores sit behind a storefront password. Asks can still crawl them: enter the password on the config dialog's Advanced tab — it's in Shopify under Online Store → Preferences — and Asks reads past the password page.
If a crawl runs into a password page it can't get past, the website row shows Password-protected store (or Store password not accepted if the saved password stopped working) with an Enter password prompt. The review picker shows the same notice and lists the store's real pages once the password is saved.
Syncing and re-crawls
Press Sync on a website row to re-crawl it, or let auto-refresh handle it on schedule. Re-syncs are incremental: every page's content is hashed, and pages whose content hasn't changed are skipped — nothing is re-indexed. Only genuinely new pages use new articles; changed pages refresh in place. Every sync also refreshes the site's link map, so pages added to your site since the last sync become selectable in the picker.
Syncs remove pages as well as add them: a healthy full crawl deletes pages that have disappeared from the site, and a selected-pages sync deletes pages you deselected. Removed pages free their article slots immediately.
Your plan sets how often each website can re-sync — weekly on Starter, daily on Business and Premium. The same frequency applies to the auto-refresh schedule and the manual Sync button (which shows when the next sync is available). A newly added website is exempt for its first 24 hours, so you can crawl, review pages, and re-sync freely while setting it up. Workspace-wide, crawls are also capped at 30 syncs per hour; beyond that, syncs are refused with "Too many website syncs in a short window" until the window passes.
Sync history (in the row's ⋯ menu) lists the last 50 syncs with status, start time, duration, pages crawled, and any error message, plus the source's success rate. If a sync goes wrong, the row says so: "Last sync incomplete" when it stopped early (for example on the article limit), "Last sync failed", or "Last sync finished with warnings".
Managing pages
Click View pages on a website to open its page list. From there you can:
- Search pages by URL or title.
- Re-crawl or delete pages individually or in bulk (up to 50 per action).
- Open a page to see exactly what was extracted — a read-only view of the synced content with its URL, extracted size, an Open original page link, a Re-crawl this page action, and Delete page. Page content is synced from the website, so it can't be edited.

Deleting a page stops the agent answering from it, but a future full sync may re-add it if it's still on the website — deselect it in Choose pages to keep it out for good. Choose pages (in the website's ⋯ menu) opens the same picker as at setup, seeded with your current selection; clearing the selection returns the source to automatic crawling.
Deleting a website
Delete in the ⋯ menu removes the source and all its crawled pages, and frees the articles they used. This can't be undone; the site can always be re-added and re-crawled later.