Your First Crawl¶
Set up your first website or social media page crawl in just a few minutes
This guide will walk you through the basic steps for starting a website crawl.
Accessing your dashboard¶
To start a crawl, you’ll need to log in using a Browsertrix account with access to crawls.
You likely have access already if:
- You signed up for a Browsertrix subscription.
- You joined an existing org and were given “crawler” permissions.
- You are the admin of a self-hosted instance.
Check if you have access by logging in. If you see a + Create New... button in the dashboard, you’re able to start a site crawl. If you don’t see this button and think that you should, contact your org administrator to update your permissions.
Choose what to archive¶
Open the webpage that you would like to archive in a new browser tab. Copy the URL; we will use this URL to start the crawl.
Choose one of the following options to tailor this guide to what you’d like to archive:
Before You Start¶
Many popular social media platforms require logging in to view the full page. To check if the social media page requires login, open the page in private browsing or incognito mode (instructions for Chrome, Safari, Edge, Firefox). If you see a prompt to sign up or login, the site will need to be provided login credentials during the crawl.
If this is the case, sign up for a new account using login credentials dedicated to archiving. Make note of these credentials in a secure place before continuing.
Can I use my existing social media account?
Although there is nothing preventing you from doing so, using your personal account to archive social media pages may put your personal login information at risk. We always recommend creating a new account dedicated to archiving to reduce the risk of accidentally archiving your private data.
Starting the crawl¶
When you log in, the first page you see is the org dashboard. If you’ve navigated away to another page, navigate back to Dashboard.
You will create a crawl workflow to set up the crawl.
-
Tap the + Create New... button and select Crawl Workflow.
-
Select the Crawl Scope that best fits what you intend to archive:
Leave the setting at Single Page.
Leave the setting at Single Page. Browsertrix automatically adjust the scope for popular social media platforms.
Choose Pages on Same Domain. This will find and archive every page that shares the same domain (e.g. your-site.com).
-
Enter the URL of the webpage that you copied earlier.
-
Configure optional settings:
Check Include directly linked pages to crawl pages linked from your target page. Including directly linked pages can improve the replay experience by preventing broken links.
If the social media page requires logging in, you can use a browser profile to provide the login credentials.
To add a browser profile:
- Scroll down and open the Browser Settings section.
- Open the Browser Profile menu and select New Browser Profile.
- Tap the Start Browser button.
- Log in using the account that you created earlier.
- Tap Create Profile to save the profile.
It can be difficult to estimate how many pages are on a website, especially if it is your first time crawling the site. Setting crawl limits is recommended as you familiarize yourself with the site crawl.
To add a crawl limit:
- Scroll down and open the Crawl Limits section.
- Set a Crawl Time Limit and/or Crawl Size Limit according to what is reasonable within your org quotas.
-
When you’re finished entering workflow settings, tap Run Crawl.
You should now see your new crawl workflow running. Give the crawler a few moments to warm up, and then watch as it crawls the webpage!
Next steps¶
After running your first crawl, you may want to:
- Include pages that the crawler didn’t visit.
- Exclude pages that shouldn’t be crawled.
- Explore all available crawl workflow setup options.
- Review the crawled content for quality assurance.
- Import previously archived content.
- Start a collection of crawled content.