HasData
Back to all posts

Airbnb Data Extraction with No Code, an API, or Python

Airbnb is an online platform for booking accommodation, allowing people to rent out or lease their houses, apartments, rooms, or other types of accommodation for short-term stays. It enables travelers to find various accommodation options in different parts of the world, including hotels, apartments, and holiday homes.

Users can browse available options, read reviews from other guests, book accommodation, and contact hosts directly through the platform. You can obtain information about properties, reviews and ratings, hosts, bookings, and availability using scrapers.

Such information about available properties can be useful not only to avid travelers and frequent business travelers but also to real estate agents and private investors who can use other properties in the area to understand what amenities and prices competitors are offering to develop their proposals.

Three routes reach that data. A no-code scraper hands you a file, an API returns the fields as JSON, and a headless browser reads the search page yourself. What follows covers all three, with the Python for the last two.

Scraping Against Calling an API

When collecting data from a website, the first option to consider is the official API. This method usually provides full data access without requiring workarounds. Although APIs may not provide all data, developers generally prefer to use them.

Scrapers are used when an official API is unavailable or impractical. In this case, a developer can create a scraper in any programming language to collect the required data automatically. This section will discuss the pros and cons of creating a Python scraper, the Airbnb API, and its alternatives.

Pros and Cons of Airbnb Scraping

Scraping Airbnb costs you differently from calling an API, and the table below puts the two side by side:

Pros of Airbnb ScrapingCons of Airbnb Scraping
1. Flexibility: Can extract any data from Airbnb listings.1. Instability: Website changes can break scraping scripts.
2. Customization: Able to tailor data extraction to specific needs.2. Maintenance: Every markup change costs you another round of selector work.
3. Independence: No reliance on Airbnb’s API availability or limitations.3. Technical Complexity: Scraping requires coding skills and maintenance.
4. Cost-effectiveness: Scraping can be cost-efficient compared to API access fees.4. Rate Limitations: Scraping may be subject to rate limitations imposed by Airbnb, affecting data retrieval speed.
5. Access to Hidden Data: Can access data not exposed through Airbnb’s API.5. Risk of Blocking: Airbnb might restrict access or IP addresses.
6. Real-time Data: Can retrieve the latest data directly from the website.6. Captcha and Bot Detection: Airbnb may deploy measures to detect scraping activities.

Weigh those against each other before deciding whether to scrape Airbnb or reach for one of the alternatives below.

The Official Airbnb API and What It Asks For

The shortest route to the data would be the official API, and Airbnb does have one. However, using it is unlikely to be successful. If you go to the official Airbnb API page, you can find the following information in the FAQ section:

The Airbnb partner page, with the answer to "How can I get access to Airbnb's API?" boxed in red

Airbnb API usage

So, the official API is not an option, but we have an alternative solution. To simplify the process of getting data from Airbnb, you can use third-party services to collect the data. If you just need a ready-made dataset, you don’t need to write any scripts. Instead, you can use our Airbnb no-code scraper.

The HasData no-code Airbnb scraper with its result limit, destination and check-in and check-out date fields

Airbnb no-code scraper

To obtain data, you only need to specify the most necessary parameters:

  1. Result Rows Limit. Enter the number of result rows you want to receive.
  2. Destination. Enter the city or town where you want to find offers.
  3. Check-in date. Enter the check-in date in the format yyyy-mm-dd.
  4. Check-out date. Enter the check-out date in the same format.

After you have specified all the necessary parameters, click the Run Scraper button to start the data collection process. When it is finished, you can download the data in JSON, CSV, or XLSX format on the right side of the screen.

Example of the data you will receive:

The spreadsheet the no-code scraper produces, one row per listing

The resulting document

This option may be more convenient if you just need to get the data. However, if you want to create your own application that constantly receives up-to-date data, we will tell you how to do it below.

Scrape Listings using Unofficial Airbnb API

Let’s start with the simpler option of using HasData’s Airbnb API to get listing data and parse it using Python.

Three things about this endpoint are worth knowing before you write the loop, because each one costs a debugging session otherwise.

Dates are required, and the error says so only if you read it. A call with a location and no dates returns HTTP 422 with {"rule": "required", "field": "checkIn"} rather than an empty result set. checkIn and checkOut both have to be there.

Price is an object, not a number. One property comes back as:

{
  "originalPrice": "$727",
  "discountedPrice": "$607",
  "qualifier": "for 4 nights",
  "breakdown": [{"description": "4 nights x $151.69", "price": "$606.74"}]
}

The strings carry currency symbols, qualifier tells you what span the total covers, and the nightly rate is inside breakdown rather than at the top. A script that treats price as a float won’t survive the first record.

The result set is not stable between calls. The same query run three times returned 18 properties each time, and only 12 to 16 of them matched the first run. Across those three identical calls the union was 24 distinct properties. Airbnb reshuffles what it shows, so page boundaries mean less than they look like they do. Deduplicate on id, and run a query more than once when coverage matters more than speed.

This API returns a JSON response in the following format:

- requestMetadata
    - id
    - status
    - url
- properties (array of objects)
    - id
    - url
    - title
    - description
    - latitude
    - longitude
    - photos (array of URLs)
    - rating
    - reviews
    - badges
    - price
      - originalPrice
      - discountedPrice
      - qualifier
      - breakdown
- pagination
    - nextPageToken
    - pageTokens (array)

One call returns 18 properties and the tokens for the pages after it. rating, reviews and badges are absent on listings that carry none, so read them with .get() rather than by subscript.

To use it, sign up and copy your personal API key:

HasData dashboard on the API Keys page, with the Default key row highlighted

API Keys under Workspace

Let’s proceed with building the scraper. If you want to get the resulting script, go to the Google Colaboratory.

You need Python 3.10 or above. This is a requirement for all the scripts we’ll be using. To install the necessary libraries, open a command prompt or terminal and run the following commands:

pip install requests pandas

Create a new project and put the whole request in one file. The dates go in as YYYY-MM-DD, and params= does the encoding, which matters the moment a location has a space in it:

import json
from datetime import date, timedelta

import requests

API_URL = "https://api.hasdata.com/scrape/airbnb/listing/"
HEADERS = {"x-api-key": "YOUR-API-KEY"}

# a fixed date turns into a past date the month after you write it, so the
# search window is built from today
search = {
    "location": "New York",
    "checkIn": (date.today() + timedelta(days=30)).isoformat(),
    "checkOut": (date.today() + timedelta(days=32)).isoformat(),
}

response = requests.get(API_URL, params=search, headers=HEADERS, timeout=60)
if response.status_code == 200:
    data = response.json()
    with open("airbnb_listing.json", "w", encoding="utf-8") as f:
        json.dump(data, f, ensure_ascii=False, indent=2)
    print(f"{len(data.get('properties', []))} listings saved to airbnb_listing.json")
else:
    print(f"Error: {response.status_code}")

Building that URL with an f-string is the one thing to avoid here. A space is safe, since requests encodes it on the way out. An ampersand in a value is not, because it starts a new parameter, and a hash turns everything after it into a fragment the server never sees. Passing the values through params lets requests encode each one.

As a result, we will obtain a JSON object:

{
    "requestMetadata": {
        "id": "f21693a0-6d10-411c-b20c-cd7c582de687",
        "status": "ok",
        "url": "https://www.airbnb.com/s/New-York/homes?checkin=2026-10-18&checkout=2026-10-20"
    },
    "properties": [
        {
            "id": "1502499830496127121",
            "url": "https://www.airbnb.com/rooms/1502499830496127121",
            "title": "Apartment in Guttenberg",
            "latitude": 40.79042,
            "longitude": -74.00366,
            "description": "Modern Condo Near NYC Skyline + Parking",
            "photos": [
                "https://a0.muscache.com/im/pictures/hosting/Hosting-1502499830496127121/original/f32f0f57-e44b-42f8-be0b-6f3d8d160c94.png",
                "https://a0.muscache.com/im/pictures/hosting/Hosting-1502499830496127121/original/b9ed9869-56bd-4714-810d-fb4aae0b7a7a.png"
            ],
            "rating": 4.97,
            "reviews": 74,
            "badges": ["Guest favorite"],
            "price": {
                "originalPrice": "$563",
                "discountedPrice": "$494",
                "qualifier": "for 2 nights",
                "breakdown": [
                    {"description": "2 nights x $281.50", "price": "$563.00"},
                    {"description": "Special offer", "price": "-$69.75"}
                ]
            }
        }
    ],
    "pagination": {
        "nextPageToken": "eyJzZWN0aW9uX29mZnNldCI6MCwiaXRlbXNfb2Zmc2V0IjoxOCwidmVyc2lvbiI6MX0=",
        "pageTokens": [
            "eyJzZWN0aW9uX29mZnNldCI6MCwiaXRlbXNfb2Zmc2V0IjowLCJ2ZXJzaW9uIjoxfQ==",
            "eyJzZWN0aW9uX29mZnNldCI6MCwiaXRlbXNfb2Zmc2V0IjoxOCwidmVyc2lvbiI6MX0="
        ]
    }
}

The API hands back the fields directly, so there are no selectors to maintain and no exit addresses to rotate, which is most of what running a scraper costs.

Scrape Airbnb Property Data with API

Search gives you a page of listings. The property endpoint gives you one listing in full, with the description, the amenity list, the host record and the coordinates. A ready-made version of this script is in Google Colaboratory.

import json

import requests

API_URL = "https://api.hasdata.com/scrape/airbnb/property/"
HEADERS = {"x-api-key": "YOUR-API-KEY"}

PROPERTIES = [
    "https://www.airbnb.com/rooms/946842435422127304",
    "https://www.airbnb.com/rooms/49662898",
]


def fetch_property(listing_url):
    # params= rather than an f-string, so a listing URL carrying its own
    # query parameters is encoded instead of breaking the request
    response = requests.get(API_URL, params={"url": listing_url},
                            headers=HEADERS, timeout=60)
    response.raise_for_status()
    return response.json()["property"]


collected = []
for listing_url in PROPERTIES:
    prop = fetch_property(listing_url)
    collected.append(prop)
    with open(f"airbnb_property_{prop['id']}.json", "w", encoding="utf-8") as f:
        json.dump(prop, f, ensure_ascii=False, indent=2)

with open("airbnb_properties.json", "w", encoding="utf-8") as f:
    json.dump(collected, f, ensure_ascii=False, indent=2)

Each property comes back with the same fourteen top-level keys, so a record from one listing lines up with a record from another: id, title, overview, description, rating, reviews, address, latitude, longitude, photos, safetyAndPropertyInfo, guestCapacity, host and amenities.

The amenities list has a trap in it. It carries an available flag, and the entries where that flag is false are things the place does not have. On the two listings above, one returned 35 amenities with 9 of them unavailable, and the other returned 15 with 8 unavailable. A membership test reports every one of those absences as a feature, and on the 35-amenity listing Wifi is among them, so a filter written that way keeps a place that has no wifi:

has_wifi = any(a["title"] == "Wifi" and a["available"] for a in prop["amenities"])

The two listings differ in size as much as in content, at 52 photos against 18 and 47 reviews against 17, so anything that writes fixed-width rows needs a plan for the ragged ones.

Scrape Airbnb Using Headless Browser

If you want to take the more challenging route and write your own scraper, you will need a headless browser. This will mimic a user’s behavior and allow you to navigate the page and collect data.

Scraping without an API is difficult because Airbnb’s content is dynamically generated. Trying to use simple requests to get data from the page won’t work. But if you want to try it on your own, we have provided a script on Google Collaboratory that you can use to try scraping the page with Requests, BeautifulSoup or Scrapy and save the resulting HTML code to a file. The result will be quite visual:

The HTML saved by a plain requests fetch, opened in a browser, where the header and search bar render but every listing card is an empty grey placeholder

The result

So the practical route is a headless browser driven by Selenium.

Installing Necessary Libraries

Selenium and Pandas are the only two libraries this scraper needs. Install them with the package manager:

pip install selenium pandas

For Selenium, you may also need to download and install the webdriver separately. However, this is not necessary in the latest versions. If you haven’t used headless browsers, you can find an article on scraping with Selenium in our blog.

Understanding Airbnb’s Website Structure

A ready-made version of this script is in Google Colaboratory. You can download and run it on your own PC since Google doesn’t allow running Web Driver in Colab Research.

A search on Airbnb with the housing parameters filled in shows what there is to extract:

An Airbnb search results grid with the first card annotated, boxing the location line, the bed count, the nightly price, the total price and the rating

Airbnb listing

So, if we extract data from the Airbnb search page, we will get the following information:

  1. Property title.
  2. Number of rooms.
  3. Price. Airbnb returns the price for one night and the specified number of days.
  4. Rating. This parameter is not available for all listings, and this should be considered during scraping.

The link is the next thing to look at. A search started from the home page produces a URL of this shape:

https://www.airbnb.com/?tab_id=home_tab&refinement_paths%5B%5D=%2Fhomes&search_mode=flex_destinations_search&flexible_trip_lengths%5B%5D=one_week&monthly_start_date=2025-04-01&monthly_length=3&monthly_end_date=2025-07-01&category_tag=Tag%3A789&price_filter_input_type=0&channel=EXPLORE&date_picker_type=calendar&checkin=2025-03-30&checkout=2025-04-02&source=structured_search_input_header&search_type=filter_change

A search that names a location produces this shape instead:

https://www.airbnb.com/s/New-York--United-States/homes?tab_id=home_tab&refinement_paths%5B%5D=%2Fhomes&flexible_trip_lengths%5B%5D=one_week&monthly_start_date=2025-04-01&monthly_length=3&monthly_end_date=2025-07-01&price_filter_input_type=0&channel=EXPLORE&date_picker_type=calendar&checkin=2025-03-30&checkout=2025-04-02&source=structured_search_input_header&search_type=autocomplete_click&price_filter_num_nights=3&query=New%20York%2C%20United%20States&place_id=ChIJOwE7_GTtwokRFq0uOwLSE9g

Let’s identify the main settings we’ll be changing later on:

  1. Location. This is where you want to search, like New-York--United-States.
  2. search_mode: It’s the search mode. “flex_destinations_search” means flexible search for destinations.
  3. checkin: check-in date, in YYYY-MM-DD form.
  4. checkout: check-out date, same form.
  5. source: Source of the search. “structured_search_input_header” indicates searching from the structured input header.

These settings are enough to create a working link for scraping.

Extracting Listing Details

Create a file with a .py extension. The whole scraper is one script, and the parts worth reading are the wait and the finally:

from datetime import date, timedelta

import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

CITY = "New-York"
COUNTRY = "United-States"
CHECKIN = (date.today() + timedelta(days=30)).isoformat()
CHECKOUT = (date.today() + timedelta(days=35)).isoformat()

CARD = '[data-testid="card-container"]'

options = Options()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)

data = []
try:
    url = (f"https://www.airbnb.com/s/{CITY}--{COUNTRY}/homes"
           f"?checkin={CHECKIN}&checkout={CHECKOUT}"
           "&source=structured_search_input_header")
    driver.get(url)

    # the page renders its cards after the first paint, so reading the DOM
    # straight after get() returns an empty list
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, CARD))
    )

    # and one card is all that exists at that moment. The rest load on scroll,
    # so the wait has to be followed by one before the grid is complete
    driver.execute_script("window.scrollBy(0, 4000)")
    WebDriverWait(driver, 10).until(
        lambda d: len(d.find_elements(By.CSS_SELECTOR, CARD)) > 1
    )

    for listing in driver.find_elements(By.CSS_SELECTOR, CARD):
        def text_of(selector):
            found = listing.find_elements(By.CSS_SELECTOR, selector)
            return found[0].text.strip() if found else None

        data.append({
            "Link": listing.find_element(By.TAG_NAME, "a").get_attribute("href"),
            "Title": text_of('[data-testid="listing-card-title"]'),
            "Name": text_of('[data-testid="listing-card-name"]'),
            "Subtitle": text_of('[data-testid="listing-card-subtitle"] > span'),
        })
finally:
    driver.quit()

print(f"{len(data)} listings")

WebDriverWait holds until a card exists, because driver.get returns as soon as the document loads and the cards arrive after that, so a script that reads the DOM immediately collects nothing. And driver.quit() goes in a finally, because a Chrome left running after an exception stays in memory until you kill it by hand, and a loop over cities leaves one behind per failure.

text_of returns None rather than raising when an element is absent. find_element throws on a missing element and takes the whole page down with it, and on a results grid one card built differently from the rest is enough to trigger that.

The price is the field to think about before you copy anything. Inside a card it arrives in elements whose class names are generated, of the span._1y74zjx kind, and those change with any redesign. The row around them has a stable hook, [data-testid="price-availability-row"], whose text is the displayed amount followed by the word total, so a scraper that wants the number rather than the breakdown can take the row and parse the text. Taking the listing URL from this script and passing it to an API that returns the price as a field is the other way out.

The Same Scraper in Playwright

Playwright does the same job with the waiting built into the locator rather than bolted on, which is the reason to reach for it here.

from datetime import date, timedelta

from playwright.sync_api import sync_playwright

CITY = "New-York"
COUNTRY = "United-States"
CHECKIN = (date.today() + timedelta(days=30)).isoformat()
CHECKOUT = (date.today() + timedelta(days=35)).isoformat()

CARD = '[data-testid="card-container"]'

data = []
with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    try:
        url = (f"https://www.airbnb.com/s/{CITY}--{COUNTRY}/homes"
               f"?checkin={CHECKIN}&checkout={CHECKOUT}"
               "&source=structured_search_input_header")
        page.goto(url)
        page.wait_for_selector(CARD, timeout=15000)

        # same story as Selenium: one card exists here, the rest need a scroll
        page.mouse.wheel(0, 4000)
        page.wait_for_function(
            "sel => document.querySelectorAll(sel).length > 1", arg=CARD
        )

        for listing in page.locator(CARD).all():
            def text_of(selector):
                found = listing.locator(selector)
                return found.first.inner_text().strip() if found.count() else None

            data.append({
                "Link": listing.locator("a").first.get_attribute("href"),
                "Title": text_of('[data-testid="listing-card-title"]'),
                "Name": text_of('[data-testid="listing-card-name"]'),
                "Subtitle": text_of('[data-testid="listing-card-subtitle"] > span'),
            })
    finally:
        browser.close()

print(f"{len(data)} listings")

get_attribute("href") returns what the attribute says, so a relative /rooms/1643343865444978960?... comes back where Selenium’s version of the same call returns the absolute URL. Join it against the site root yourself if you need a link you’d want to store.

locator.count() is a query against the live page rather than a snapshot, which is why text_of above checks it before reading. Selenium’s find_elements returns a list that is already fixed at the moment you asked.

Neither driver’s faster than the other on this page. Both spend the same wait on the same render, and the scroll is what decides whether you collect one listing or eighteen.

Save Airbnb Data

Pandas turns that list of dictionaries into a table and writes it out in whichever format the next step wants:

df = pd.DataFrame(data)
df.to_csv(f"airbnb_{CITY}_{COUNTRY}_listings.csv", index=False)
df.to_json(f"airbnb_{CITY}_{COUNTRY}_listings.json", orient="records")
df.to_excel(f"airbnb_{CITY}_{COUNTRY}_listings.xlsx", index=False)

The CSV opens like this, one row per card:

Headless browser scraping result, a CSV of Airbnb listings

The resulting CSV file

The first two lines run on the pandas install from earlier. to_excel does not, because the Excel writer is an optional extra, so pip install openpyxl comes first or that line raises an ImportError about a missing optional dependency.

What Blocks an Airbnb Scraper

Airbnb renders its listing cards after the initial HTML arrives, which is the first thing that stops a scraper and the easiest to mistake for a block. A plain requests.get() returns HTTP 200 and a page with every card as an empty placeholder, so nothing raises and nothing’s parsed.

The second thing is the address. Datacenter ranges are scored differently from residential ones, and a few hundred requests from one of them is enough for the site to start answering with a challenge instead of a page. Residential exits and a slower rate are the usual answer, and a rotating pool matters more than any single exit does.

The third is the CAPTCHA that follows. It arrives as a page with HTTP 200, the same as the placeholder case, so a scraper that only checks the status code cannot tell the two apart. Check for a field you’d expect on a real listing rather than for a status.

The order to try them is the same as the order they appear. Run a browser first, because the rendering problem is the one that hits everybody. Add proxies when volume, not the page, is what breaks. Treat a CAPTCHA as a signal that the first two were not enough rather than as something to solve on its own.

Airbnb Scraping Challenges

When scraping Airbnb, review their Terms of Service and make sure your use targets only publicly available data and complies with applicable law. A scraping API can help you collect public data reliably without managing the infrastructure yourself.

Another challenge you will face is updating the website structure periodically. If Airbnb changes the structure of its website, scraping may become unusable, as your scripts may stop working due to changes in the HTML or CSS markup. In such cases, you should update your scripts to match the new website structure.

Here are some additional points to consider:

  1. Use a scraping API. It carries the proxy pool and the rendering, so a markup change costs you a selector rather than a rebuild.
  2. Use a proxy. Routing requests through rotating proxies spreads them across IP addresses for more reliable collection of public data.

Conclusion and Takeaways

With the data you gather from Airbnb through web scraping, you can analyze the real estate market more effectively. You can create your own price comparison services or recommendation systems and use this data in your marketing strategies and advertisements. For example, businesses and entrepreneurs can use this data for targeted marketing and advertising campaigns aimed at audiences interested in renting properties.

If you prefer not to gather this data using Python, other programming languages are suitable for web scraping. For instance, JavaScript (using the Puppeteer library), Ruby (using the Nokogiri library), and Go (using the Colly library) can be good alternatives. Additionally, you can use a ready-made Airbnb no-code scraper, which provides a user interface for scraping data without coding.

Valentina Skakun
Valentina Skakun
Valentina is a software engineer who builds data extraction tools before writing about them. With a strong background in Python, she also leverages her experience in JavaScript, PHP, R, and Ruby to reverse-engineer complex web architectures.If data renders in a browser, she will find a way to script its extraction.
Articles

Might Be Interesting