Adding website URLs | IcloneU
Knowledge Library

Adding website URLs

Point IcloneU's crawler at a single starting URL and it brings the site's pages, linked files and embedded media into your Knowledge library. Here's the whole flow — the scope controls, the discover-and-review step, and what to do when the site changes.

4 min read Beginner Updated July 6, 2026

Crawling your website is the quickest way to teach a Clone everything you've already published. You give IcloneU one starting URL; the crawler discovers the site's internal links on its own, extracts the page text, downloads linked files and transcribes embedded media — and the result lands in your Knowledge library as a single searchable source.

Before you begin

You'll need an existing Knowledge library (you can create one on the way), the full URL of the site or page you want to crawl, and the right to ingest that content — the platform doesn't check copyright or terms-of-service compliance for you. New to libraries? Start with What is the Knowledge Library?

What the crawler collects

A website source brings in three kinds of content — and the dialog you'll meet below has a switch for the last two:

  • Page content — the text of every page the crawler discovers from your starting URL.
  • Linked files — documents the crawled pages link to are downloaded and indexed alongside the pages themselves.
  • Embedded media — audio and video found on the pages is transcribed, with YouTube, Vimeo, Dailymotion, Facebook, TikTok, SoundCloud, Spotify, iVoox, podcast feeds and HLS/DASH streams all recognised.

After a whole YouTube channel or playlist rather than a website? That's the other web-family source: choose YouTube instead of Website when you add the source, and every video is imported and transcribed individually.

Add a website to your library

Configuration takes under a minute; the crawl itself runs in the background, so you can keep working while it does.

  1. Open your library's sources

    Select Knowledge in the left sidebar. On the Sources sub-tab, find the target library and click the chevron at the left of its row to expand it. No library yet? Click + New library — one is created immediately, named "New Library".

  2. Choose Website from the Add Knowledge dropdown

    Click the Add Knowledge dropdown and choose Website. (The same dropdown is also home to the other source types — YouTube, Files & Folders and Link Existing Collection.)

  3. Enter the starting URL

    Paste the full URL, including https://. One starting point is enough — the crawler discovers the site's internal links from there. If the URL is malformed, the dialog shows validation feedback.

  4. Set the crawl scope, then Discover

    Give the collection a name and adjust the scope controls — each one is explained in the table below. Then click Discover to find pages before anything is ingested.

    Tip

    The defaults suit most sites: Include Files and Include Media on, with Max per channel preset to 100 if you switch Limit Media on. Often the only thing worth typing is the Collection Name.

  5. Review what was found, then Process

    After discovery, the dialog lists everything it found, each item with a checkbox, with Pages / Files / Media filter buttons to view one category at a time. Untick what you don't want, then click Process (N). Items are queued and processed in the background — the source row's status runs Queued → Review Items → Ready, and the Activity sub-tab logs the progress of each phase. While discovery runs, the source row shows a live status like Discovering: 22 found, 140 reviewedfound counts unique pages and often plateaus while the crawler works through pagination and duplicates; the reviewed number keeps climbing the whole time, so a rising reviewed count means the crawl is healthy even when found stops moving.

The Add Website dialog: URL field, scope controls and the Discover button
One dialog covers the whole flow: URL, scope, Discover, review, Process.

When the row shows Ready, everything the crawl extracted is searchable by any Clone attached to the library.

The scope controls, explained

Step 4's controls decide how far the crawl reaches and how much it brings back. These are the controls and their defaults:

ControlDefaultWhat it's for
Collection NameThe label this source shows in your library
Include FilesOnDownload the files that crawled pages link to
Include MediaOnTranscribe embedded audio and video found on pages
Limit MediaOffCap how much embedded media the crawl takes on
Max per channel100 (range 1–2000)Upper bound on items collected per channel
Excluded SubdomainsComma-separated subdomains to skip, e.g. music, audio, store
Excluded PathsComma-separated paths to skip, e.g. /legal, /privacy, /admin
Tip

Legal boilerplate, admin areas and store pages rarely make good knowledge. A minute spent on Excluded Paths and Excluded Subdomains keeps the crawl lean — and the answers cleaner.

When the website changes

A website source is a snapshot taken at the moment of the crawl, not a live mirror. Your Clone reads from that indexed snapshot — publish a price change or a new blog post five minutes after the crawl and the Clone won't see it until the source is re-synchronised. Three ways to bring it up to date:

  • Scheduled re-sync — on the library's Synchronization sub-tab, set a re-crawl cadence and the platform re-fetches the site on that schedule.
  • Manual re-sync — same sub-tab; trigger a one-off re-crawl when you've pushed a change you want reflected straight away.
  • Replace the source — delete the website source and add it again. A fresh crawl from scratch is sometimes cleaner than a re-sync when the site's structure has changed significantly.

Two things a re-sync does not do: pages that are no longer reachable from your starting URL are moved to the Recycle Bin on the next sync — reported on the Activity sub-tab as Removed (N) — no longer on the site and restorable from the Recycle Bin as a set — and content that has moved behind a login simply stops appearing, because the crawler doesn't authenticate. If a specific page must survive regardless, ingest it separately as a file source (export it to PDF, say, and upload that).

Note

After every re-sync, the Activity sub-tab logs the run — worth a spot-check that it finished and the new content is in. Cadences, cost guards and the rest of the sync configuration are covered in Keeping your library up to date.

Frequently asked

Setting one up takes under a minute; the crawl and processing then run in the background. How long that takes depends on the site's size and depth and on how much embedded media has to be downloaded and transcribed — small text-heavy sites are typically quick, media-heavy ones take longer. Treat those as rough guides rather than guaranteed times (IcloneU doesn't publish processing SLAs), and watch the Activity sub-tab for live progress. During discovery the source row also shows a live status like Discovering: 22 found, 140 reviewed — a climbing reviewed count means the crawl is still working, even when found has stopped moving.

The site is probably blocking crawlers — via robots.txt or anti-bot protection. The crawler can't get past that; the workaround is a different ingestion route, such as exporting the pages (to PDF, say) and uploading them as file sources — see Adding documents and files.

The usual cause is login-gated content: the crawler isn't known to authenticate against arbitrary sites, so pages that require signing in won't be captured. Export those pages manually and upload them as file sources instead.

Pages that vanish from the site aren't deleted outright. A re-sync moves them to the Recycle Bin and lists them on the Activity sub-tab as Removed (N) — no longer on the site. Open the Recycle Bin to restore the removed set.

The hosting platform may have changed its access policy, or the media may be geo-restricted. If the content matters, try adding the media directly via its own URL, where a source type accepts it.

There's a better route: choose YouTube instead of Website in the Add Knowledge dropdown, give it a channel or playlist URL, and every video is imported and transcribed. (YouTube videos embedded in pages you crawl are transcribed as part of the website source too.)

Was this guide helpful?
Thanks for the feedback!

Last updated July 6, 2026 · Knowledge Library

Reconnecting to the server… Reload
🗙
Connecting…
Connection lost
Reconnecting to the server…
We couldn't reconnect automatically.