A login page stops a scraper in one of five ways, and each one takes a different fix:
| Obstacle | How you handle it | When you meet it |
|---|---|---|
| Plain login form | requests.Session() with the posted credentials | Simple sites, admin panels, internal dashboards |
| CSRF token | Read the hidden token from the page, send it with the form | Most modern form frameworks (Laravel, Django, Rails) |
| WAF or bot check | A real browser through SeleniumBase UC mode | Cloudflare and similar in front of the login |
| CAPTCHA at login | A solving service or a human step, or an official API instead | reCAPTCHA and hCaptcha on the form |
| Two-factor auth | No automated login, reuse a manual session or use an API token | Accounts with SMS or authenticator codes required |
The sections below run from the simplest case to the hardest.
Prerequisites
You’ll need Python 3.11 or newer. The code hasn’t been tested on older versions, so it might not work with them.
The examples use requests and BeautifulSoup for the simple cases, Selenium where the page needs a browser, and Scrapy for one of them:
pip install requests beautifulsoup4 selenium scrapyAny library that supports a web driver works in place of Selenium.
Basic Authentication
Basic authentication comes in two shapes, a login form that takes a POST and an endpoint that takes the credentials in the request itself.
Sending Data to a Website
The simplest case is basic authentication through a login form, where verifying your identity means sending your username and password in a POST request, no tokens or extra protection involved:

Here’s an example of a simple form that just takes a username and password as POST parameters. If you open DevTools and check the Network tab during login, you’ll see that those are the only values being sent:

For a site like this, requests alone is enough to log in and scrape the protected data:
import requests
login_url = "https://www.scrapingcourse.com/login"
payload = {
'email': 'admin@example.com',
'password': 'password'
}
response = requests.post(login_url, data=payload)
print(response.status_code)But this will only get you through the login once. Most of the time, you’ll need to stay logged in while making multiple requests to scrape data from different pages.
To do that, you can use a session. Let’s update the script to store the login state in a session:
import requests
session = requests.Session()
login_url = "https://www.scrapingcourse.com/login"
response = session.get(login_url)
response.raise_for_status()
payload = {
'email': 'admin@example.com',
'password': 'password'
}
post_response = session.post(login_url, data=payload)
post_response.raise_for_status()Now you can keep using that session to access other pages that require authentication:
protected_page = session.get("https://www.scrapingcourse.com/dashboard")Sites with this little protection are rare, and when you find one, BeautifulSoup handles the parsing.
from bs4 import BeautifulSoupAfter you make the request, you can extract the data you need using CSS selectors:
soup = BeautifulSoup(protected_page.text, 'html.parser')
for product in soup.select(".product-item"):
name = product.select_one(".product-name")
price = product.select_one(".product-price")
link_tag = product.select_one("a")
image_tag = product.select_one("img")
product_name = name.text.strip() if name else "N/A"
product_price = price.text.strip() if price else "N/A"
product_url = link_tag['href'] if link_tag else "N/A"
product_image = image_tag['src'] if image_tag else "N/A"
print(f"Name: {product_name}")
print(f"Price: {product_price}")
print(f"Product URL: {product_url}")
print(f"Image URL: {product_image}")
print("-" * 100)Here’s the result:

Full code example:
import requests
from bs4 import BeautifulSoup
session = requests.Session()
login_url = "https://www.scrapingcourse.com/login"
response = session.get(login_url)
response.raise_for_status()
payload = {
'email': 'admin@example.com',
'password': 'password'
}
post_response = session.post(login_url, data=payload)
post_response.raise_for_status()
protected_page = session.get("https://www.scrapingcourse.com/dashboard")
soup = BeautifulSoup(protected_page.text, 'html.parser')
for product in soup.select(".product-item"):
name = product.select_one(".product-name")
price = product.select_one(".product-price")
link_tag = product.select_one("a")
image_tag = product.select_one("img")
product_name = name.text.strip() if name else "N/A"
product_price = price.text.strip() if price else "N/A"
product_url = link_tag['href'] if link_tag else "N/A"
product_image = image_tag['src'] if image_tag else "N/A"
print(f"Name: {product_name}")
print(f"Price: {product_price}")
print(f"Product URL: {product_url}")
print(f"Image URL: {product_image}")
print("-" * 100)Basic Auth is old, and it still runs plenty of API backends and internal dashboards behind HTTPS. Any HTTP client speaks it, curl and requests included.
Using an API Endpoint
Let’s look at a more realistic example of basic authentication, specifically, making a request to an endpoint that expects auth credentials.

The endpoint takes the credentials in the URL path, so you can set your own and watch the header work:
import requests
# httpbin echoes back whether the credentials matched
url = "https://httpbin.org/basic-auth/name/pass"
response = requests.get(url, auth=("name", "pass"))
print("Status Code:", response.status_code)
print("Response JSON:", response.json())auth= builds the Authorization: Basic header for you, base64 and all. Sending the wrong pair returns 401 with an empty body, which is the quickest way to tell a credential problem from a parsing one.
The response looks like this:

Now let’s see what happens if we send the wrong credentials:
bad_response = requests.get(url, auth=("name", "wrongpass"))
print("Status Code (failed):", bad_response.status_code) A failed auth attempt returns an empty body, so there is nothing to print the way the successful response had. Instead, we’ll just show the status code:

Here’s the full code for the example:
import requests
url = "https://httpbin.org/basic-auth/name/pass"
response = requests.get(url, auth=("name", "pass"))
print("Status Code:", response.status_code)
print("Response JSON:", response.json())
bad_response = requests.get(url, auth=("name", "wrongpass"))
print("Status Code (failed):", bad_response.status_code) This is the typical form of basic authentication you’ll most likely run into.
CSRF Token Authentication
In more complex authentication cases, the server generates a random token (like a CSRF token) that you need to send along with your login and password. The server checks this token with every request, and if it’s missing or incorrect, the request gets rejected.

This will still be a POST request with the same parameters as in the previous example, but now we’ll include the token. Open DevTools and go to the Network tab, just like before:

You can usually find the token on the page, inside a hidden input field:

So the example stays basically the same. The only difference is that before sending the request, we need to grab the token from the page and pass it as one of the parameters:
soup = BeautifulSoup(response.text, 'html.parser')
csrf_token = soup.select_one("input[name='_token']")['value'] if soup.select_one("input[name='_token']") else None
payload = {
'email': 'admin@example.com',
'password': 'password',
'_token': csrf_token
}Other than that, the code doesn’t change. We’ll still get the data we need in the response:

Here’s the full code:
import requests
from bs4 import BeautifulSoup
session = requests.Session()
login_url = "https://www.scrapingcourse.com/login"
response = session.get(login_url)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
csrf_token = soup.select_one("input[name='_token']")['value'] if soup.select_one("input[name='_token']") else None
payload = {
'email': 'admin@example.com',
'password': 'password',
'_token': csrf_token
}
post_response = session.post(login_url, data=payload)
post_response.raise_for_status()
protected_page = session.get("https://www.scrapingcourse.com/dashboard")
protected_page.raise_for_status()
soup = BeautifulSoup(protected_page.text, 'html.parser')
for product in soup.select(".product-item"):
name = product.select_one(".product-name")
price = product.select_one(".product-price")
link_tag = product.select_one("a")
image_tag = product.select_one("img")
product_name = name.text.strip() if name else "N/A"
product_price = price.text.strip() if price else "N/A"
product_url = link_tag['href'] if link_tag else "N/A"
product_image = image_tag['src'] if image_tag else "N/A"
print(f"Name: {product_name}")
print(f"Price: {product_price}")
print(f"Product URL: {product_url}")
print(f"Image URL: {product_image}")
print("-" * 100)I recommend this dynamic approach, always extract the token at runtime instead of hardcoding it. Tokens can change, and a hardcoded value will eventually break.
WAF (Web Application Firewall) Authentication
A WAF (Web Application Firewall) sits in front of the login and runs its own checks. Those come in five shapes:
- Two-Factor Authentication (2FA). After entering a username and password, you also need to enter a code sent to your device (via SMS or an authenticator app).
- CAPTCHA during login. The server asks the user to prove they’re not a bot by solving a CAPTCHA challenge.
- JavaScript Challenges. Some sites, especially behind Cloudflare, require the client to run JS to prove it’s a real browser, not a bot.
- OAuth or OpenID Connect. In this case, the server redirects the user to an external service (like Google or Facebook) for authentication. After that, the server returns a token used for further requests.
- IP blocking and geofiltering. Some sites block requests from suspicious or unauthorized IPs or entire regions. If your IP doesn’t match the allowed range, requests get denied.
Some of these protections can be bypassed (e.g., IP filters via proxies). Others, like OAuth, OpenID Connect, or 2FA, can’t really be bypassed, you just have to go through them.
Two-factor authentication is the hard stop for automation. When a site requires a one-time code from an SMS, an authenticator app, or a hardware key, no scraper can supply it, because the whole point of the second factor is that only the account holder can. Where 2FA is mandatory, automated login is not possible, and the realistic routes are a session you established manually and then reuse (see the cookie and storage_state sections below) or an official API with a long-lived token.
Other mechanisms only trigger if the site thinks you’re acting like a bot. You can avoid those by mimicking real user behavior.
Three Python libraries strip the automation signals a plain driver leaks:
pip install undetected-chromedriver seleniumbase selenium-stealthundetected-chromedriver is a modified ChromeDriver that hides the automation signals a plain one leaks. seleniumbase in UC mode does the same and adds its own handling for challenge pages, and builds on the first. selenium-stealth adjusts the browser fingerprint instead, so it patches what the page can read rather than what the driver announces.
Pick what fits your project. We’ll use the second option listed above: SeleniumBase in UC mode.
Logging In Behind a WAF with SeleniumBase
uc=True starts SeleniumBase in undetected mode. uc_open_with_reconnect then opens the page with the driver detached for the seconds the challenge needs to settle, and uc_gui_click_captcha clicks the checkbox through the operating system pointer instead of through JavaScript:
from seleniumbase import SB
login_url = "https://www.scrapingcourse.com/login/cf-antibot"
email = "admin@example.com"
password = "password"
with SB(uc=True) as sb:
sb.uc_open_with_reconnect(login_url, reconnect_time=6)
sb.uc_gui_click_captcha()The rest of the logic stays close to the previous examples, but now runs through Selenium:
sb.type('input[name="email"]', email)
sb.type('input[name="password"]', password)
sb.click('button[type="submit"]')
sb.wait_for_element(".product-item", timeout=10)
products = sb.find_elements("div.product-item")
for product in products:
name = product.text.split('\n')[0] if product.text else "N/A"
price = product.find_element("css selector", ".product-price").text if product.find_elements("css selector", ".product-price") else "N/A"
product_url = product.find_element("tag name", "a").get_attribute('href') if product.find_elements("tag name", "a") else "N/A"
image_url = product.find_element("tag name", "img").get_attribute('src') if product.find_elements("tag name", "img") else "N/A"
print(f"Name: {name}")
print(f"Price: {price}")
print(f"Product URL: {product_url}")
print(f"Image URL: {image_url}")
print("-" * 100)The output holds the same product fields the requests version returned.

Here’s the full code:
from seleniumbase import SB
login_url = "https://www.scrapingcourse.com/login/cf-antibot"
email = "admin@example.com"
password = "password"
with SB(uc=True) as sb:
sb.uc_open_with_reconnect(login_url, reconnect_time=6)
sb.uc_gui_click_captcha()
sb.type('input[name="email"]', email)
sb.type('input[name="password"]', password)
sb.click('button[type="submit"]')
sb.wait_for_element(".product-item", timeout=10)
products = sb.find_elements("div.product-item")
for product in products:
name = product.text.split('\n')[0] if product.text else "N/A"
price = product.find_element("css selector", ".product-price").text if product.find_elements("css selector", ".product-price") else "N/A"
product_url = product.find_element("tag name", "a").get_attribute('href') if product.find_elements("tag name", "a") else "N/A"
image_url = product.find_element("tag name", "img").get_attribute('src') if product.find_elements("tag name", "img") else "N/A"
print(f"Name: {name}")
print(f"Price: {price}")
print(f"Product URL: {product_url}")
print(f"Image URL: {image_url}")
print("-" * 100)The browser does what a visitor’s browser does, so the page treats the run as an ordinary session.
Login on a Real Site
Let’s take a real website instead of a test one and adapt our previous code for HasData.

The login block will mostly stay the same, only the input and button field selectors will change:
from seleniumbase import SB
from selenium.webdriver.common.by import By
with SB(uc=True) as sb:
sb.open("https://app.hasdata.com/sign-in")
sb.type('input[name="email"]', "YOUR-EMAIL")
sb.type('input[name="password"]', "YOUR-PASSWORD\n") # the trailing newline submits the formAfter that, you can go to any page you need (for example, the marketplace page) and scrape the data.
First, figure out which tags contain the data you want, like this:
sb.wait_for_element("div.items-center", timeout=10)
sb.click('a[href="/marketplace"]')
sb.wait_for_element("ul.place-content-stretch > li", timeout=10)
reviews = sb.find_elements("css selector", "ul.place-content-stretch > li")
for r in reviews:
divs = r.find_elements(By.CSS_SELECTOR, "div")
print(divs[0].text.strip().split('\n')[0])Here’s how it works in practice:
So, you’ll get a list of available no-code scrapers:

The same data is available without running a browser yourself through HasData’s no-code scrapers and scraping APIs, which take the request, render the page on their side and return the HTML or parsed JSON.
Authentication with reCAPTCHA v2
A CAPTCHA on the form is where automated login usually stops. The demo below asks for two fields, a CAPTCHA and a button click, which is close to a real flow.

There are a few ways to deal with CAPTCHAs:
- Use CAPTCHA-solving services. You send them the site key and URL, and they return a token after solving it (usually by a human or a bot). You then insert the token into the form.
- Reuse the token. Sometimes the token you get after solving a CAPTCHA can be reused multiple times.
- Look for workarounds. For example, you can use a browser extension or an API from the target service to get the needed data without triggering a CAPTCHA at all.
Some CAPTCHAs never show themselves, reCAPTCHA v2 invisible among them, and those score the session instead of asking a question.
Let’s look at an example using a browser extension to solve reCAPTCHA v2 and the SeleniumBase library to mask the bot.
First, download any CAPTCHA-solving extension from the Chrome Web Store in .crx format. Convert it to .zip, unzip it, and set the path to the folder in your script:
from seleniumbase import SB
extension_path = "captcha"
# extension_dir loads the unpacked extension when the browser starts
with SB(uc=True, extension_dir=extension_path) as sb:
sb.open("https://recaptcha-demo.appspot.com/recaptcha-v2-checkbox.php")
sb.type('input[name="ex-a"]', "my_username")
sb.type('input[name="ex-b"]', "my_password")
sb.click('button[type="submit"]')extension_dir points at a directory rather than an archive, so unpack the extension on disk first and give it that folder.
The script will look like this when running:
Full code:
from seleniumbase import SB
extension_path = "captcha"
with SB(uc=True, extension_dir=extension_path) as sb:
sb.open("https://recaptcha-demo.appspot.com/recaptcha-v2-checkbox.php")
sb.type('input[name="ex-a"]', "my_username")
sb.type('input[name="ex-b"]', "my_password")
sb.click('button[type="submit"]')Extensions cover the common reCAPTCHA v2 checkbox and little beyond it. For the rest, a solving API is the route that actually returns a token.
Using Scrapy’s FormRequest for Authentication
FormRequest.from_response reads the form out of the response you hand it, carries every hidden field across, and keeps the action URL and the method the form declares, so a CSRF token you never named goes along with the post. A hand-rolled FormRequest with a formdata dict sends only what you list, so it works on a login page today and breaks the week the site adds a field.
Two things go wrong with it in practice. When a page holds more than one form the search takes the first, so a site with a newsletter box above the login sends your credentials to the wrong place, and formid or formxpath fixes that. And when JavaScript draws the form, there is nothing in the response to read, so from_response raises instead of posting, which is the point to switch to a browser.
The code below is the spider itself, with the project setup the Scrapy basics guide covers left out.
First, let’s set some basic parameters like the spider name and the target URL:
import scrapy
from scrapy.http import FormRequest
class TestLoginSpider(scrapy.Spider):
name = 'test_login'
start_urls = ['URL']
# Here will be your codeThen the login goes out through FormRequest.from_response:
def parse(self, response):
return FormRequest.from_response(
response,
formdata={
'login': 'username',
'password': 'password'
},
callback=self.after_login
)After that, define what happens after the form is submitted, for example, check if the login was successful:
def after_login(self, response):
if "username" in response.text:
self.logger.info("Login successful")
else:
self.logger.error("Login failed")
return
yield scrapy.Request(
url='new_URL',
callback=self.parse_protected
)If login fails, we log an error. If it works, we move on to a protected page:
def parse_protected(self, response):
self.logger.info("Now on protected page")That’s where you can add your scraping logic.
Here’s the full code:
import scrapy
from scrapy.http import FormRequest
class TestLoginSpider(scrapy.Spider):
name = 'test_login'
start_urls = ['URL']
def parse(self, response):
return FormRequest.from_response(
response,
formdata={
'login': 'username',
'password': 'password'
},
callback=self.after_login
)
def after_login(self, response):
if "username" in response.text:
self.logger.info("Login successful")
else:
self.logger.error("Login failed")
return
yield scrapy.Request(
url='new_URL',
callback=self.parse_protected
)
def parse_protected(self, response):
self.logger.info("Now on protected page")Running the spider logs Login successful when the credentials land, and Scrapy’s cookie middleware carries the session into every request that follows, so parse_protected opens on an authenticated page with no further setup.
Cookie Reuse
To avoid logging in every time, you can reuse cookies, just keep in mind that they usually don’t live long.
With requests and a session, pickle writes the jar to disk and reads it back:
import pickle
# after a successful login
with open("cookies.pkl", "wb") as f:
pickle.dump(session.cookies, f)
# in the next run, instead of logging in again
with open("cookies.pkl", "rb") as f:
session.cookies.update(pickle.load(f))Selenium keeps its own jar, so the same idea goes through the driver object, which is sb.driver under SeleniumBase:
cookies = driver.get_cookies()
# later, on a page from the same domain
for cookie in cookies:
driver.add_cookie(cookie)add_cookie throws if the browser is not already on a page from that domain, which is the usual reason a restored session appears to do nothing. Check the expiry as well, since a jar that loads without error can still be entirely stale.
Saving a Whole Session with Playwright storage_state
The pickle and Selenium approaches above save cookies by hand. Playwright has this built in through storage_state, which writes cookies and localStorage together to one JSON file and restores the whole session from it, so later runs never touch the login form. Log in once and save the state:
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://the-internet.herokuapp.com/login")
page.fill("#username", "tomsmith")
page.fill("#password", "SuperSecretPassword!")
page.click("button[type='submit']")
page.wait_for_selector(".flash.success")
page.context.storage_state(path="auth_state.json")
browser.close()Every later run loads that file into a fresh context and lands in the authenticated area directly:
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(storage_state="auth_state.json")
page = context.new_page()
page.goto("https://the-internet.herokuapp.com/secure")
print(page.inner_text("h2")) # Secure Area
browser.close()An unauthenticated visit to that secure page redirects back to the login form, so printing the Secure Area heading confirms the saved session carried the login across runs. The state file expires with its cookies just like the pickle jar, so refresh it when a restored run lands back on the login page.
Conclusion
If an API exists for the data, take it. It is the least code to write and the least to maintain. Without one, start with a session, add the hidden token when the form carries one, and move to a browser only when the page runs a check requests cannot answer. Two-factor auth is where the automated path ends.
For authentication, you’ll likely need libraries like SeleniumBase, UndetectedBrowser, and others to hide the fact that you’re using code to scrape data.


