HasData
Back to all posts

How to Scrape a Website that Requires Login with Python

A login page stops a scraper in one of five ways, and each one takes a different fix:

ObstacleHow you handle itWhen you meet it
Plain login formrequests.Session() with the posted credentialsSimple sites, admin panels, internal dashboards
CSRF tokenRead the hidden token from the page, send it with the formMost modern form frameworks (Laravel, Django, Rails)
WAF or bot checkA real browser through SeleniumBase UC modeCloudflare and similar in front of the login
CAPTCHA at loginA solving service or a human step, or an official API insteadreCAPTCHA and hCaptcha on the form
Two-factor authNo automated login, reuse a manual session or use an API tokenAccounts with SMS or authenticator codes required

The sections below run from the simplest case to the hardest.

Prerequisites

You’ll need Python 3.11 or newer. The code hasn’t been tested on older versions, so it might not work with them.

The examples use requests and BeautifulSoup for the simple cases, Selenium where the page needs a browser, and Scrapy for one of them:

pip install requests beautifulsoup4 selenium scrapy

Any library that supports a web driver works in place of Selenium.

Basic Authentication

Basic authentication comes in two shapes, a login form that takes a POST and an endpoint that takes the credentials in the request itself.

Sending Data to a Website

The simplest case is basic authentication through a login form, where verifying your identity means sending your username and password in a POST request, no tokens or extra protection involved:

The scrapingcourse Login demo form, with Email Address and Password fields and the demo credentials in a banner

Here’s an example of a simple form that just takes a username and password as POST parameters. If you open DevTools and check the Network tab during login, you’ll see that those are the only values being sent:

DevTools Network panel, Payload tab, showing the login POST carrying only email and password as form data

For a site like this, requests alone is enough to log in and scrape the protected data:

import requests


login_url = "https://www.scrapingcourse.com/login"


payload = {
    'email': 'admin@example.com',
    'password': 'password'
}


response = requests.post(login_url, data=payload)


print(response.status_code)

But this will only get you through the login once. Most of the time, you’ll need to stay logged in while making multiple requests to scrape data from different pages.

To do that, you can use a session. Let’s update the script to store the login state in a session:

import requests


session = requests.Session()


login_url = "https://www.scrapingcourse.com/login"


response = session.get(login_url)
response.raise_for_status()


payload = {
    'email': 'admin@example.com',
    'password': 'password'
}


post_response = session.post(login_url, data=payload)
post_response.raise_for_status()

Now you can keep using that session to access other pages that require authentication:

protected_page = session.get("https://www.scrapingcourse.com/dashboard")

Sites with this little protection are rare, and when you find one, BeautifulSoup handles the parsing.

from bs4 import BeautifulSoup

After you make the request, you can extract the data you need using CSS selectors:

soup = BeautifulSoup(protected_page.text, 'html.parser')


for product in soup.select(".product-item"):
    name = product.select_one(".product-name")
    price = product.select_one(".product-price")
    link_tag = product.select_one("a")
    image_tag = product.select_one("img")


    product_name = name.text.strip() if name else "N/A"
    product_price = price.text.strip() if price else "N/A"
    product_url = link_tag['href'] if link_tag else "N/A"
    product_image = image_tag['src'] if image_tag else "N/A"


    print(f"Name: {product_name}")
    print(f"Price: {product_price}")
    print(f"Product URL: {product_url}")
    print(f"Image URL: {product_image}")
    print("-" * 100)

Here’s the result:

Terminal output listing the scraped products, each with a name, a price and product and image URLs

Full code example:

import requests
from bs4 import BeautifulSoup


session = requests.Session()


login_url = "https://www.scrapingcourse.com/login"


response = session.get(login_url)
response.raise_for_status()


payload = {
    'email': 'admin@example.com',
    'password': 'password'
}


post_response = session.post(login_url, data=payload)
post_response.raise_for_status()


protected_page = session.get("https://www.scrapingcourse.com/dashboard")


soup = BeautifulSoup(protected_page.text, 'html.parser')


for product in soup.select(".product-item"):
    name = product.select_one(".product-name")
    price = product.select_one(".product-price")
    link_tag = product.select_one("a")
    image_tag = product.select_one("img")


    product_name = name.text.strip() if name else "N/A"
    product_price = price.text.strip() if price else "N/A"
    product_url = link_tag['href'] if link_tag else "N/A"
    product_image = image_tag['src'] if image_tag else "N/A"


    print(f"Name: {product_name}")
    print(f"Price: {product_price}")
    print(f"Product URL: {product_url}")
    print(f"Image URL: {product_image}")
    print("-" * 100)

Basic Auth is old, and it still runs plenty of API backends and internal dashboards behind HTTPS. Any HTTP client speaks it, curl and requests included.

Using an API Endpoint

Let’s look at a more realistic example of basic authentication, specifically, making a request to an endpoint that expects auth credentials.

The browser's native sign-in dialog for httpbin.org asking for a username and password

The endpoint takes the credentials in the URL path, so you can set your own and watch the header work:

import requests

# httpbin echoes back whether the credentials matched
url = "https://httpbin.org/basic-auth/name/pass"

response = requests.get(url, auth=("name", "pass"))
print("Status Code:", response.status_code)
print("Response JSON:", response.json())

auth= builds the Authorization: Basic header for you, base64 and all. Sending the wrong pair returns 401 with an empty body, which is the quickest way to tell a credential problem from a parsing one.

The response looks like this:

Terminal output reading Status Code 200 and Response JSON with authenticated set to True

Now let’s see what happens if we send the wrong credentials:

bad_response = requests.get(url, auth=("name", "wrongpass"))
print("Status Code (failed):", bad_response.status_code) 

A failed auth attempt returns an empty body, so there is nothing to print the way the successful response had. Instead, we’ll just show the status code:

Terminal output from the failed attempt, showing status code 401

Here’s the full code for the example:

import requests

url = "https://httpbin.org/basic-auth/name/pass" 

response = requests.get(url, auth=("name", "pass"))
print("Status Code:", response.status_code)               
print("Response JSON:", response.json())                  

bad_response = requests.get(url, auth=("name", "wrongpass"))
print("Status Code (failed):", bad_response.status_code) 

This is the typical form of basic authentication you’ll most likely run into.

CSRF Token Authentication

In more complex authentication cases, the server generates a random token (like a CSRF token) that you need to send along with your login and password. The server checks this token with every request, and if it’s missing or incorrect, the request gets rejected.

The Login with CSRF demo form, with Email Address and Password fields and the demo credentials in a banner

This will still be a POST request with the same parameters as in the previous example, but now we’ll include the token. Open DevTools and go to the Network tab, just like before:

DevTools Network panel, Payload tab, with the form data showing _token highlighted above email and password

You can usually find the token on the page, inside a hidden input field:

DevTools Elements panel with the hidden input named _token highlighted inside the login form markup

So the example stays basically the same. The only difference is that before sending the request, we need to grab the token from the page and pass it as one of the parameters:

soup = BeautifulSoup(response.text, 'html.parser')
csrf_token = soup.select_one("input[name='_token']")['value'] if soup.select_one("input[name='_token']") else None


payload = {
    'email': 'admin@example.com',
    'password': 'password',
    '_token': csrf_token  
}

Other than that, the code doesn’t change. We’ll still get the data we need in the response:

Terminal output listing the products scraped after the login that carried the CSRF token

Here’s the full code:

import requests
from bs4 import BeautifulSoup

session = requests.Session()

login_url = "https://www.scrapingcourse.com/login"

response = session.get(login_url)
response.raise_for_status()

soup = BeautifulSoup(response.text, 'html.parser')
csrf_token = soup.select_one("input[name='_token']")['value'] if soup.select_one("input[name='_token']") else None

payload = {
    'email': 'admin@example.com',
    'password': 'password',
    '_token': csrf_token  
}

post_response = session.post(login_url, data=payload)
post_response.raise_for_status()

protected_page = session.get("https://www.scrapingcourse.com/dashboard")
protected_page.raise_for_status()

soup = BeautifulSoup(protected_page.text, 'html.parser')

for product in soup.select(".product-item"):
    name = product.select_one(".product-name")
    price = product.select_one(".product-price")
    link_tag = product.select_one("a")
    image_tag = product.select_one("img")

    product_name = name.text.strip() if name else "N/A"
    product_price = price.text.strip() if price else "N/A"
    product_url = link_tag['href'] if link_tag else "N/A"
    product_image = image_tag['src'] if image_tag else "N/A"

    print(f"Name: {product_name}")
    print(f"Price: {product_price}")
    print(f"Product URL: {product_url}")
    print(f"Image URL: {product_image}")
    print("-" * 100)

I recommend this dynamic approach, always extract the token at runtime instead of hardcoding it. Tokens can change, and a hardcoded value will eventually break.

WAF (Web Application Firewall) Authentication

A WAF (Web Application Firewall) sits in front of the login and runs its own checks. Those come in five shapes:

  1. Two-Factor Authentication (2FA). After entering a username and password, you also need to enter a code sent to your device (via SMS or an authenticator app).
  2. CAPTCHA during login. The server asks the user to prove they’re not a bot by solving a CAPTCHA challenge.
  3. JavaScript Challenges. Some sites, especially behind Cloudflare, require the client to run JS to prove it’s a real browser, not a bot.
  4. OAuth or OpenID Connect. In this case, the server redirects the user to an external service (like Google or Facebook) for authentication. After that, the server returns a token used for further requests.
  5. IP blocking and geofiltering. Some sites block requests from suspicious or unauthorized IPs or entire regions. If your IP doesn’t match the allowed range, requests get denied.

Some of these protections can be bypassed (e.g., IP filters via proxies). Others, like OAuth, OpenID Connect, or 2FA, can’t really be bypassed, you just have to go through them.

Two-factor authentication is the hard stop for automation. When a site requires a one-time code from an SMS, an authenticator app, or a hardware key, no scraper can supply it, because the whole point of the second factor is that only the account holder can. Where 2FA is mandatory, automated login is not possible, and the realistic routes are a session you established manually and then reuse (see the cookie and storage_state sections below) or an official API with a long-lived token.

Other mechanisms only trigger if the site thinks you’re acting like a bot. You can avoid those by mimicking real user behavior.

Three Python libraries strip the automation signals a plain driver leaks:

pip install undetected-chromedriver seleniumbase selenium-stealth

undetected-chromedriver is a modified ChromeDriver that hides the automation signals a plain one leaks. seleniumbase in UC mode does the same and adds its own handling for challenge pages, and builds on the first. selenium-stealth adjusts the browser fingerprint instead, so it patches what the page can read rather than what the driver announces.

Pick what fits your project. We’ll use the second option listed above: SeleniumBase in UC mode.

Logging In Behind a WAF with SeleniumBase

uc=True starts SeleniumBase in undetected mode. uc_open_with_reconnect then opens the page with the driver detached for the seconds the challenge needs to settle, and uc_gui_click_captcha clicks the checkbox through the operating system pointer instead of through JavaScript:

from seleniumbase import SB

login_url = "https://www.scrapingcourse.com/login/cf-antibot"

email = "admin@example.com"
password = "password"

with SB(uc=True) as sb:
    sb.uc_open_with_reconnect(login_url, reconnect_time=6)
    sb.uc_gui_click_captcha()

The rest of the logic stays close to the previous examples, but now runs through Selenium:

    sb.type('input[name="email"]', email)
    sb.type('input[name="password"]', password)
    sb.click('button[type="submit"]')

    sb.wait_for_element(".product-item", timeout=10)

    products = sb.find_elements("div.product-item")

    for product in products:
        name = product.text.split('\n')[0] if product.text else "N/A"
        price = product.find_element("css selector", ".product-price").text if product.find_elements("css selector", ".product-price") else "N/A"
        product_url = product.find_element("tag name", "a").get_attribute('href') if product.find_elements("tag name", "a") else "N/A"
        image_url = product.find_element("tag name", "img").get_attribute('src') if product.find_elements("tag name", "img") else "N/A"

        print(f"Name: {name}")
        print(f"Price: {price}")
        print(f"Product URL: {product_url}")
        print(f"Image URL: {image_url}")
        print("-" * 100)

The output holds the same product fields the requests version returned.

Terminal output from the SeleniumBase run, listing the same product names, prices and URLs

Here’s the full code:

from seleniumbase import SB

login_url = "https://www.scrapingcourse.com/login/cf-antibot"

email = "admin@example.com"
password = "password"

with SB(uc=True) as sb:
    sb.uc_open_with_reconnect(login_url, reconnect_time=6)
    sb.uc_gui_click_captcha()

    sb.type('input[name="email"]', email)
    sb.type('input[name="password"]', password)
    sb.click('button[type="submit"]')

    sb.wait_for_element(".product-item", timeout=10)

    products = sb.find_elements("div.product-item")

    for product in products:
        name = product.text.split('\n')[0] if product.text else "N/A"
        price = product.find_element("css selector", ".product-price").text if product.find_elements("css selector", ".product-price") else "N/A"
        product_url = product.find_element("tag name", "a").get_attribute('href') if product.find_elements("tag name", "a") else "N/A"
        image_url = product.find_element("tag name", "img").get_attribute('src') if product.find_elements("tag name", "img") else "N/A"

        print(f"Name: {name}")
        print(f"Price: {price}")
        print(f"Product URL: {product_url}")
        print(f"Image URL: {image_url}")
        print("-" * 100)

The browser does what a visitor’s browser does, so the page treats the run as an ordinary session.

Login on a Real Site

Let’s take a real website instead of a test one and adapt our previous code for HasData.

HasData sign-in form with email and password fields, GitHub and Google sign-in options, and the Continue button

The login block will mostly stay the same, only the input and button field selectors will change:

from seleniumbase import SB
from selenium.webdriver.common.by import By


with SB(uc=True) as sb:
    sb.open("https://app.hasdata.com/sign-in")
    sb.type('input[name="email"]', "YOUR-EMAIL")
    sb.type('input[name="password"]', "YOUR-PASSWORD\n")  # the trailing newline submits the form

After that, you can go to any page you need (for example, the marketplace page) and scrape the data.

First, figure out which tags contain the data you want, like this:

    sb.wait_for_element("div.items-center", timeout=10)
    sb.click('a[href="/marketplace"]')

    sb.wait_for_element("ul.place-content-stretch > li", timeout=10)

    reviews = sb.find_elements("css selector", "ul.place-content-stretch > li")
    for r in reviews:
        divs = r.find_elements(By.CSS_SELECTOR, "div")
        print(divs[0].text.strip().split('\n')[0])

Here’s how it works in practice:

So, you’ll get a list of available no-code scrapers:

Terminal output listing HasData's no-code scrapers, from Google Search Results Scraper to Crunchbase Scraper

The same data is available without running a browser yourself through HasData’s no-code scrapers and scraping APIs, which take the request, render the page on their side and return the HTML or parsed JSON.

Authentication with reCAPTCHA v2

A CAPTCHA on the form is where automated login usually stops. The demo below asks for two fields, a CAPTCHA and a button click, which is close to a real flow.

The reCAPTCHA v2 demo form, with two example inputs, an I'm not a robot checkbox and a Submit button

There are a few ways to deal with CAPTCHAs:

  1. Use CAPTCHA-solving services. You send them the site key and URL, and they return a token after solving it (usually by a human or a bot). You then insert the token into the form.
  2. Reuse the token. Sometimes the token you get after solving a CAPTCHA can be reused multiple times.
  3. Look for workarounds. For example, you can use a browser extension or an API from the target service to get the needed data without triggering a CAPTCHA at all.

Some CAPTCHAs never show themselves, reCAPTCHA v2 invisible among them, and those score the session instead of asking a question.

Let’s look at an example using a browser extension to solve reCAPTCHA v2 and the SeleniumBase library to mask the bot.

First, download any CAPTCHA-solving extension from the Chrome Web Store in .crx format. Convert it to .zip, unzip it, and set the path to the folder in your script:

from seleniumbase import SB

extension_path = "captcha"

# extension_dir loads the unpacked extension when the browser starts
with SB(uc=True, extension_dir=extension_path) as sb:
    sb.open("https://recaptcha-demo.appspot.com/recaptcha-v2-checkbox.php")
    sb.type('input[name="ex-a"]', "my_username")
    sb.type('input[name="ex-b"]', "my_password")
    sb.click('button[type="submit"]')

extension_dir points at a directory rather than an archive, so unpack the extension on disk first and give it that folder.

The script will look like this when running:

Full code:

from seleniumbase import SB

extension_path = "captcha"  

with SB(uc=True, extension_dir=extension_path) as sb:
    sb.open("https://recaptcha-demo.appspot.com/recaptcha-v2-checkbox.php")  
    sb.type('input[name="ex-a"]', "my_username")
    sb.type('input[name="ex-b"]', "my_password")
    sb.click('button[type="submit"]')

Extensions cover the common reCAPTCHA v2 checkbox and little beyond it. For the rest, a solving API is the route that actually returns a token.

Using Scrapy’s FormRequest for Authentication

FormRequest.from_response reads the form out of the response you hand it, carries every hidden field across, and keeps the action URL and the method the form declares, so a CSRF token you never named goes along with the post. A hand-rolled FormRequest with a formdata dict sends only what you list, so it works on a login page today and breaks the week the site adds a field.

Two things go wrong with it in practice. When a page holds more than one form the search takes the first, so a site with a newsletter box above the login sends your credentials to the wrong place, and formid or formxpath fixes that. And when JavaScript draws the form, there is nothing in the response to read, so from_response raises instead of posting, which is the point to switch to a browser.

The code below is the spider itself, with the project setup the Scrapy basics guide covers left out.

First, let’s set some basic parameters like the spider name and the target URL:

import scrapy
from scrapy.http import FormRequest


class TestLoginSpider(scrapy.Spider):
    name = 'test_login'
    start_urls = ['URL']


    # Here will be your code

Then the login goes out through FormRequest.from_response:

    def parse(self, response):
        return FormRequest.from_response(
            response,
            formdata={
                'login': 'username',
                'password': 'password'
            },
            callback=self.after_login
        )

After that, define what happens after the form is submitted, for example, check if the login was successful:

    def after_login(self, response):
        if "username" in response.text: 
            self.logger.info("Login successful")
        else:
            self.logger.error("Login failed")
            return


        yield scrapy.Request(
            url='new_URL',
            callback=self.parse_protected
        )

If login fails, we log an error. If it works, we move on to a protected page:

    def parse_protected(self, response):
        self.logger.info("Now on protected page")

That’s where you can add your scraping logic.

Here’s the full code:

import scrapy
from scrapy.http import FormRequest


class TestLoginSpider(scrapy.Spider):
    name = 'test_login'
    start_urls = ['URL']


    def parse(self, response):
        return FormRequest.from_response(
            response,
            formdata={
                'login': 'username',
                'password': 'password'
            },
            callback=self.after_login
        )


    def after_login(self, response):
        if "username" in response.text: 
            self.logger.info("Login successful")
        else:
            self.logger.error("Login failed")
            return


        yield scrapy.Request(
            url='new_URL',
            callback=self.parse_protected
        )


    def parse_protected(self, response):
        self.logger.info("Now on protected page")

Running the spider logs Login successful when the credentials land, and Scrapy’s cookie middleware carries the session into every request that follows, so parse_protected opens on an authenticated page with no further setup.

To avoid logging in every time, you can reuse cookies, just keep in mind that they usually don’t live long.

With requests and a session, pickle writes the jar to disk and reads it back:

import pickle

# after a successful login
with open("cookies.pkl", "wb") as f:
    pickle.dump(session.cookies, f)

# in the next run, instead of logging in again
with open("cookies.pkl", "rb") as f:
    session.cookies.update(pickle.load(f))

Selenium keeps its own jar, so the same idea goes through the driver object, which is sb.driver under SeleniumBase:

cookies = driver.get_cookies()

# later, on a page from the same domain
for cookie in cookies:
    driver.add_cookie(cookie)

add_cookie throws if the browser is not already on a page from that domain, which is the usual reason a restored session appears to do nothing. Check the expiry as well, since a jar that loads without error can still be entirely stale.

Saving a Whole Session with Playwright storage_state

The pickle and Selenium approaches above save cookies by hand. Playwright has this built in through storage_state, which writes cookies and localStorage together to one JSON file and restores the whole session from it, so later runs never touch the login form. Log in once and save the state:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://the-internet.herokuapp.com/login")
    page.fill("#username", "tomsmith")
    page.fill("#password", "SuperSecretPassword!")
    page.click("button[type='submit']")
    page.wait_for_selector(".flash.success")
    page.context.storage_state(path="auth_state.json")
    browser.close()

Every later run loads that file into a fresh context and lands in the authenticated area directly:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context(storage_state="auth_state.json")
    page = context.new_page()
    page.goto("https://the-internet.herokuapp.com/secure")
    print(page.inner_text("h2"))   # Secure Area
    browser.close()

An unauthenticated visit to that secure page redirects back to the login form, so printing the Secure Area heading confirms the saved session carried the login across runs. The state file expires with its cookies just like the pickle jar, so refresh it when a restored run lands back on the login page.

Conclusion

If an API exists for the data, take it. It is the least code to write and the least to maintain. Without one, start with a session, add the hidden token when the form carries one, and move to a browser only when the page runs a check requests cannot answer. Two-factor auth is where the automated path ends.

For authentication, you’ll likely need libraries like SeleniumBase, UndetectedBrowser, and others to hide the fact that you’re using code to scrape data.

Valentina Skakun
Valentina Skakun
Valentina is a software engineer who builds data extraction tools before writing about them. With a strong background in Python, she also leverages her experience in JavaScript, PHP, R, and Ruby to reverse-engineer complex web architectures.If data renders in a browser, she will find a way to script its extraction.
Articles

Might Be Interesting