TikTok is a dynamic, JavaScript-heavy platform. Most of the data you see in the browser isn’t in the page’s initial HTML. A request to the page returns partial information, with no video list, no profile stats and sometimes not even a username.
Complete data lives in the scripts and the JSON endpoints the page calls as it loads, rather than in the markup.
In this guide, we’ll walk you through several ways to extract profile data, video lists, comments, and search results. We’ll also explain how to handle common issues such as IP limits, expired links, and 403 errors.
TikTok Page Structure
The first idea that comes to mind when scraping a TikTok page is to find CSS selectors or XPath paths and extract data from them:

But that data is incomplete, and extracting it this way is extremely time-consuming. There’s an easier, faster approach: in the HTML code, you’ll find a script with the ID __UNIVERSAL_DATA_FOR_REHYDRATION__ that contains all profile data in a structured JSON format.

This works for profile data and individual video pages. For other pages, it’s often simpler to find the right endpoint that returns the required data:

Those four numbers are what every later section builds on, so it is worth checking the handle resolves before writing anything else.
Scraping TikTok Profiles
Using standard parsing from the profile page, you can only extract part of the data:

Instead, we can use the script content, which contains much more detailed, actionable information:
| Field | Description |
|---|---|
username | Public handle |
nickname | Display name |
biography | Bio text, which is where a contact address appears when there is one |
verified | Whether the account carries the badge |
createTime | When the account was created |
avatarUrl, avatarThumbUrl | Profile picture, full size and thumbnail |
followers, follows, friends | Follower counts |
likes | Total likes received |
videos | Number of videos posted |
id, secUid | The two identifiers other endpoints take |
The JSON includes extra metadata you can ignore. Apply filtering logic to extract only relevant fields before saving results.
Install and import the required libraries:
Requestsfor fetching the page contentBeautifulSoupfor parsing the HTML and finding the scriptJSONfor processing the data
Before running the script, update the target URL. TikTok rate-limits unauthenticated requests to roughly 100 per hour, so higher-volume collection needs request pacing or an API. The following example uses the HasData Web Scraping API, which manages proxies, rendering, and retries for you. Provide the target URL and it returns the page content:
import json
import requests
# Get your API key at https://hasdata.com/sign-up
API_KEY = "YOUR-API-KEY"
response = requests.get(
"https://api.hasdata.com/scrape/tiktok/profile/",
params={"handle": "funny_funny66066"},
headers={"x-api-key": API_KEY},
timeout=60,
)
response.raise_for_status()
profile = response.json()["profile"]
with open("tiktok_user_detail.json", "w", encoding="utf-8") as f:
json.dump(profile, f, ensure_ascii=False, indent=4)The endpoint returns the profile already parsed, so there is no rendered page to dig through:
{
"id": "7664638705177150477",
"username": "nasa",
"nickname": "NASA",
"biography": "Making the seemingly impossible, possible.",
"verified": true,
"language": "en",
"createTime": "2026-07-20T15:55:49.000Z",
"avatarUrl": "https://p19-common-sign.tiktokcdn-us.com/...",
"followers": 1642179,
"follows": 23,
"likes": 8410893,
"videos": 42,
"friends": 17
}Contact details, when a profile carries them, sit in the biography field. This field usually surfaces the profile description, so you’ll need to use regular expressions to isolate the email address from the rest of the text:
import json
import re
# Load JSON from file
with open("tiktok_user_detail.json", "r", encoding="utf-8") as f:
data = json.load(f)
# Extract signature text
signature = data.get("biography", "")
# Find email with regex
email = re.search(r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}", signature)Keep in mind that users rarely include contact details, so don’t rely on this field being present.
Email addresses from a profile are personal data. Collect and store them only where you have a lawful basis under GDPR or CCPA, such as consent or legitimate interest.
Scraping TikTok Videos
This part is a bit more complex, since the video list isn’t included in the same JSON we parsed earlier. Still, we can use that data to get videos from the profile.
TikTok requires a signature on api/post/item_list, and the parameters below satisfy it today rather than permanently. If you would rather not track that, the TikTok Scraper API takes a handle and returns a page of videos with a nextPageToken to walk the account back.
Get the video list
To get the video list, you can use either a) a browser automation library that scrolls through the page like a real user:

Or b) an endpoint that returns all needed data:
https://www.tiktok.com/api/post/item_list/Using an endpoint involves two main steps:
- Get the
secUidvalue from the JSON we extracted in the previous section. - Use it to call the endpoint and retrieve the videos.
Note that the API returns no more than 35 videos per request. To get more, use the cursor parameter (the position or token of the last video) to load the next page.
You can check the hasMore parameter, if it’s true, there are more videos available.
Finally, extract and save the clean, actionable data from the returned JSON, excluding metadata and unnecessary clutter.
import requests
import json
from bs4 import BeautifulSoup
import time
profile_url = "https://www.tiktok.com/@funny_funny66066" # put the username page here
api_base = "https://www.tiktok.com/api/post/item_list/"
headers = {
"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36",
"accept": "*/*",
"referer": profile_url,
}
# fetch profile page and extract secUid
r = requests.get(profile_url, headers=headers, timeout=15)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
script = soup.find("script", id="__UNIVERSAL_DATA_FOR_REHYDRATION__")
if not script or not script.string:
raise SystemExit("Could not find rehydration script on profile page.")
raw = json.loads(script.string)
secuid = (
raw.get("__DEFAULT_SCOPE__", {})
.get("webapp.user-detail", {})
.get("userInfo", {})
.get("user", {})
.get("secUid")
)
if not secuid:
raise SystemExit("secUid not found in page JSON. Page structure may differ.")
print("secUid:", secuid)
# paginated fetch using secUid
result = []
cursor = 0
has_more = True
while has_more:
params = {
"aid": "1988",
"count": "35",
"cursor": cursor,
"device_platform": "web_pc",
"secUid": secuid,
}
resp = requests.get(api_base, headers=headers, params=params, timeout=15)
# handle non-json responses
try:
data = resp.json()
except ValueError:
print("Non-JSON response, status:", resp.status_code)
break
for item in data.get("itemList", []):
# collect only the clean video shape you asked for
video = {
"profileURL": profile_url,
"id": item.get("id"),
"desc": item.get("desc"),
"createTime": item.get("createTime"),
"duration": item.get("video", {}).get("duration"),
"cover": item.get("video", {}).get("cover"),
"originCover": item.get("video", {}).get("originCover"),
"dynamicCover": item.get("video", {}).get("dynamicCover"),
"playAddr": item.get("video", {}).get("playAddr"),
"downloadAddr": item.get("video", {}).get("downloadAddr"),
"videoID": item.get("video", {}).get("videoID"),
"width": item.get("video", {}).get("width"),
"height": item.get("video", {}).get("height"),
"size": item.get("video", {}).get("size"),
"definition": item.get("video", {}).get("definition"),
"ratio": item.get("video", {}).get("ratio"),
"videoQuality": item.get("video", {}).get("videoQuality"),
"stats": {
"playCount": item.get("stats", {}).get("playCount"),
"diggCount": item.get("stats", {}).get("diggCount"),
"commentCount": item.get("stats", {}).get("commentCount"),
"shareCount": item.get("stats", {}).get("shareCount"),
"collectCount": item.get("stats", {}).get("collectCount")
},
"music": {
"authorName": item.get("music", {}).get("authorName"),
"title": item.get("music", {}).get("title"),
"playUrl": item.get("music", {}).get("playUrl"),
"duration": item.get("music", {}).get("duration")
},
"textExtra": item.get("textExtra", []),
"challenges": [ch.get("title") for ch in item.get("challenges", [])]
}
result.append(video)
# pagination tokens
cursor = data.get("cursor", 0)
has_more = bool(data.get("hasMore", False))
time.sleep(0.5)
print("Collected videos:", len(result))
# Save cleaned output
with open("videos_clean_full.json", "w", encoding="utf-8") as f:
json.dump(result, f, ensure_ascii=False, indent=2)The list is now on disk, which matters because the download links in it go stale faster than the metadata does.
Download TikTok Videos
One of the fields in the API response (and in our saved JSON) is the download link, downloadAddr.
We could download videos directly during data extraction, but to keep things organized, we’ll do that in a separate script.
We’ll use the links from the saved JSON file. The downloadAddr links in TikTok’s API responses expire within an hour or so, and an old one returns a 403.
import requests
import os
# Folder for downloaded videos
download_folder = "tiktok_videos"
os.makedirs(download_folder, exist_ok=True)
# Download a single video from URL
def download_video(url, filename):
try:
# Stream download to avoid memory overload
headers = {
"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36",
"accept": "*/*",
"referer": "https://www.tiktok.com",
}
response = requests.get(url, headers=headers, stream=True)
response.raise_for_status()
filepath = os.path.join(download_folder, filename)
with open(filepath, "wb") as f:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
print(f"[OK] {filename} downloaded")
except Exception as e:
print(f"[ERROR] {filename} failed: {e}")
# Load JSON with video info
with open("videos_clean_full.json", "r", encoding="utf-8") as f:
videos = json.load(f)
# Iterate over videos and download
for video in videos:
download_url = video.get("downloadAddr")
video_id = video.get("id")
if download_url:
filename = f"{video_id}.mp4"
download_video(download_url, filename)
else:
print(f"[WARN] No download link for video {video_id}")The same pagination token drives every later page, so store it with the results rather than in a variable that dies with the process.
Scraping TikTok Comments
The next step is extracting comments from videos.
As with the video list, the comments endpoint takes a videoId and returns the comments with author, likes and reply counts, paginated the same way.

Use TikTok’s comment API endpoint to extract comments. TikTok stops paginating at 5,000 comments per video:
import requests
import json
# list of TikTok video IDs to fetch comments from
video_ids = ["7518505042194861325", "7501766490077678891"]
# number of comments to fetch per video
count = 30
# standard desktop browser User-Agent
headers = {
"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/152.0.0.0 Safari/537.36"
}
comments = []
# iterate through each video ID and fetch comments
for vid in video_ids:
url = f"https://www.tiktok.com/api/comment/list/?aid=1988&aweme_id={vid}&count={count}"
resp = requests.get(url, headers=headers)
# attempt to parse the JSON response
try:
data = resp.json()
except ValueError:
print(f"Non-JSON response for video {vid}, status:", resp.status_code)
continue
# loop through each comment object
for c in data.get("comments", []):
comments.append(
{
"video_id": c.get("aweme_id"),
"video_desc": c["share_info"].get("desc"),
"video_title": c["share_info"].get("title"),
"video_url": c["share_info"].get("url"),
"comment_id": c.get("cid"),
"text": c.get("text"),
"likes": c.get("digg_count"),
"language": c.get("comment_language"),
"timestamp": c.get("create_time"),
"author": {
"name": c["user"].get("nickname"),
"username": c["user"].get("unique_id"),
"user_id": c["user"].get("uid"),
"avatar": (
c["user"]["avatar_thumb"]["url_list"][0]
if c["user"].get("avatar_thumb")
else None
),
},
"comment_url": c["share_info"].get("url"),
}
)
# save all collected comments into a JSON file
with open("comments.json", "w", encoding="utf-8") as f:
json.dump(comments, f, ensure_ascii=False, indent=2)
print(f"Saved {len(comments)} comments to comments.json")Comments arrive newest first, so a run that stops early keeps the recent ones and loses the thread history.
Scraping TikTok Search
Search scraping is the hardest part, because TikTok dynamically generates headers. Endpoints like:
The same API takes a keyword and a type of video or user, which is the same split the two internal URLs below make.
api/search/general/full/or:
api/search/item/full/require valid, time-sensitive headers.
To capture the responses, use a headless browser (e.g., Playwright) and intercept API calls directly.
import json
from playwright.sync_api import sync_playwright
keyword = "cat funny video"
search_url = f"https://www.tiktok.com/search?q={keyword}"
def extract_useful_data(api_response):
# Extract only the fields we care about from TikTok API
useful = []
data_list = api_response.get("data", [])
for entry in data_list:
item = entry.get("item", {})
video = item.get("video", {})
author = item.get("author", {})
music = item.get("music", {})
stats = item.get("stats", {})
hashtags = [
tag.get("hashtagName")
for tag in item.get("textExtra", [])
if tag.get("hashtagName")
]
# Build simplified dict with relevant info
useful.append({
"id": item.get("id"),
"desc": item.get("desc"),
"createTime": item.get("createTime"),
"video_url": video.get("playAddr"),
"cover": video.get("cover"),
"author_name": author.get("nickname"),
"author_id": author.get("id"),
"author_uniqueId": author.get("uniqueId"),
"author_avatar": author.get("avatarThumb"),
"music_title": music.get("title"),
"music_author": music.get("authorName"),
"hashtags": hashtags,
"stats": {
"likes": stats.get("diggCount"),
"comments": stats.get("commentCount"),
"shares": stats.get("shareCount"),
"plays": stats.get("playCount"),
"favorites": stats.get("collectCount")
}
})
return useful
def run(playwright):
# Launch browser with visible window for debugging
browser = playwright.chromium.launch(headless=False)
context = browser.new_context()
page = context.new_page()
api_response = None
# Listen for all network responses and capture TikTok search API
def handle_response(response):
nonlocal api_response
if "api/search/general/full/" in response.url: # check for API endpoint
try:
api_response = response.json() # store JSON for later parsing
except Exception:
pass # ignore invalid responses
page.on("response", handle_response)
# Open TikTok search page
page.goto(search_url, timeout=60000)
page.wait_for_timeout(10000) # wait a bit for the API to respond
if api_response:
# Parse API response into simplified structure
useful_data = extract_useful_data(api_response)
# Save to JSON file
with open("search_clean.json", "w", encoding="utf-8") as f:
json.dump(useful_data, f, ensure_ascii=False, indent=2)
# Close browser cleanly
browser.close()
# Start Playwright and run scraping
with sync_playwright() as playwright:
run(playwright)If you prefer to stick with the Requests library instead of using Playwright, you can copy the headers manually from your browser and reuse them for a while, until they expire.
Libraries that wrap this for you
Two Python libraries cover the same ground, and they are not in the same condition.
| Library | Covers | State |
|---|---|---|
TikTokApi | profiles, videos, comments, search, trending | 6,600 stars, released this year, commits within the last month |
tiktokapipy | profiles, videos, challenges | 237 stars, no release in almost three years, no commits in over two |
TikTokApi drives a real browser through Playwright to get the signature TikTok wants, so it needs a browser installed and it is slower than a plain request. That is the trade it makes for not having to track the signing scheme yourself.
tiktokapipy still installs and its examples still read well, which is what makes it easy to pick by mistake. Check when it last shipped before you build on it.
Final Thoughts
Now you have a working setup for scraping TikTok profiles and related data. You can improve it by adding error handling and pacing requests to stay within TikTok’s rate limits.
The gist is, most data is already accessible via JSON scripts and endpoints, so scraping it mostly comes down to collecting it reliably or integrating it with other tools.
All examples are available in the repository.


