# Datapika: full text Site: https://datapika.com Map: https://datapika.com/llms.txt Agent instructions: https://datapika.com/AGENTS.md # Actors ## Job Board Scraper Canonical: https://datapika.com/actors/job-board-scraper Run at: https://apify.com/openclawai/job-board-scraper Price: $0.005 per job Updated: 2026-08-29 One search across eight job boards, deduplicated automatically Runs one search across eight job boards, including LinkedIn, Indeed, Glassdoor, ZipRecruiter, Naukri, and Bayt, and returns a single deduplicated dataset. Every listing carries title, company, location, posting date, job type, remote flag, direct application URL, salary minimum and maximum with currency, and company details such as industry, employee count, and rating. - Searches eight boards in one run and removes duplicate listings automatically. - Salary ranges with currency and annual normalization, plus company industry, size, and rating. - Filters for remote work, job type, posting age, distance radius, and easy apply. ### FAQ Q: How much does it cost to scrape job postings from LinkedIn and Indeed? A: This scraper charges $0.005 per job delivered, so 1,000 listings cost about $5. It searches LinkedIn, Indeed, Glassdoor, ZipRecruiter, Naukri, and Bayt in one run and only charges for results it returns. Q: Can I scrape multiple job boards at once and remove duplicates? A: Yes, this actor queries up to eight boards in a single search and deduplicates listings that appear on more than one site. Each run returns up to 100 results per board, with salary fields, company data, and direct application links on every listing. ## LinkedIn Jobs Scraper Canonical: https://datapika.com/actors/linkedin-jobs-scraper Run at: https://apify.com/openclawai/linkedin-jobs-scraper Price: $0.0005 per job Updated: 2026-08-29 Scrape LinkedIn jobs with salaries and direct apply URLs, no login required Searches LinkedIn jobs by keyword, location, company, posted date, job type, remote, and Easy Apply, returning each listing with title, company, location, salary minimum and maximum with currency and interval, job level, job function, company industry, posting date, and job URL. Optional full-description mode adds the complete description in Markdown or HTML plus the direct apply URL, and up to five keyword searches merge into one deduplicated dataset. - Filters for keyword, location, company IDs, posted-within hours, job type, remote, and Easy Apply. - Salary ranges with currency, interval, and optional annual normalization, plus full descriptions and direct apply URLs. - Up to five search terms per run, merged and deduplicated, with no LinkedIn login, cookie, or API key. ### FAQ Q: How much does it cost to scrape LinkedIn job listings? A: This scraper charges $0.0005 per job delivered, so 1,000 listings cost about $0.50 and 10,000 about $5, and you are never billed for failed requests, retries, or empty searches. A scheduled price change raises this to $0.005 per job ($5 per 1,000) from September 11, 2026. A spending cap on the run truncates output instead of overspending. Q: Do I need a LinkedIn account or API key to scrape jobs? A: No. The scraper reads public job listings the same way a logged-out browser does, so there is no LinkedIn login, cookie, or API key involved and your own account is never at risk. You only need an Apify account, and the actor can be called from the REST API, Python and JavaScript clients, or as an MCP tool for AI agents. ## Reddit Scraper Canonical: https://datapika.com/actors/reddit-scraper Run at: https://apify.com/openclawai/reddit-scraper Price: $0.002 per post Updated: 2026-08-29 Scrape Reddit posts, comments, subreddits, and AI answers on demand Pulls public Reddit data across six actions: subreddit scraping, global or in-subreddit search, comment search, subreddit discovery, single post fetch, and Reddit AI Answers. Every row includes title, body, author, upvote and comment counts, timestamps, media URLs, and optional full nested comment trees, priced from $2 per 1,000 posts. - Six actions: subreddit scraping, search, comment search, subreddit discovery, post fetch, AI answers. - Full nested comment trees with up to 20 parallel workers, no API key or login. - Pay per result from $2 per 1,000 posts, with a 100 percent success rate. ### FAQ Q: How much does it cost to scrape Reddit posts? A: This scraper charges per result, starting at $2 per 1,000 posts and $3 per 1,000 posts with full comment trees. There is no subscription and no API key, you pay only for the rows you receive. Q: Can I scrape Reddit comments without the official API? A: Yes, this actor fetches complete nested comment trees from public Reddit pages with no API key or login, using up to 20 parallel workers. It holds a 4.9 rating from 188 users and reports a 100 percent success rate. ## TikTok & Douyin Scraper Canonical: https://datapika.com/actors/tiktok-douyin-bilibili-scraper Run at: https://apify.com/openclawai/tiktok-douyin-bilibili-scraper Price: $0.001 per video Updated: 2026-08-29 Scrape TikTok, Douyin, and Bilibili videos, profiles, and comments Scrapes TikTok, Douyin, and Bilibili in one run, returning titles, descriptions, play, like, comment, and share counts, author profiles with follower counts, hashtags, music credits, and no-watermark MP4 links where the platform provides them. It also covers comments, live streams, and trending feeds through a single unified schema, up to 5,000 items per run with no API key required. - Full video records: title, stats, author, music, hashtags, and no-watermark MP4 links. - Covers videos, profiles, comments, live streams, and trending feeds, no API key required. - Flat $0.001 per video, comment, or profile, and failed URLs are never billed. ### FAQ Q: How much does it cost to scrape TikTok, Douyin, or Bilibili videos? A: Every video, comment, or profile costs $0.001, which works out to $1.00 per 1,000 videos. URLs that fail return an error row instead of data and are not billed. Q: Can I download TikTok and Douyin videos without a watermark? A: Yes, each video record includes a no-watermark MP4 link whenever the platform exposes one, across TikTok, Douyin, and Bilibili. Douyin's CDN currently rate-limits some of these direct links, so a portion of Douyin MP4 URLs can return a 403 until the platform rotates them. ## Airbnb Scraper Canonical: https://datapika.com/actors/airbnb-scraper Run at: https://apify.com/openclawai/airbnb-scraper Price: $0.001 per listing Updated: 2026-08-29 Scrape Airbnb listings, prices, reviews, calendars, and host data Pulls Airbnb search results, full room details with stay quotes, guest reviews, calendar availability, host portfolios, and Experiences from a pasted search URL or a map bounding box. Every record is a flat JSON row with a record_type field, ready for spreadsheets, BI tools, and AI pipelines. - Six modes cover listings, room details, reviews, calendars, host portfolios, and Experiences. - Search by pasted Airbnb URL or map bounding box with date, price, and amenity filters. - Pay per delivered record from $0.001, with empty and failed runs never charged. ### FAQ Q: How much does it cost to scrape Airbnb listings? A: This scraper charges a flat $0.001 per delivered record, whether it is a listing from search, a full room detail, a review, a calendar month, a host portfolio item, or an Experience, so 1,000 records of any type cost $1.00. Empty results and failed runs are not charged. Q: Can I get Airbnb prices for specific dates? A: Yes, pass checkIn and checkOut dates in YYYY-MM-DD format and the scraper returns nightly prices and full stay quotes for those dates. Calendar mode also returns month by month availability and pricing at $0.001 per calendar month, so a 12-month calendar costs $0.012. ## Google Maps Scraper Canonical: https://datapika.com/actors/google-maps-scraper Run at: https://apify.com/openclawai/google-maps-scraper Price: $0.003 per place Updated: 2026-08-29 Scrape Google Maps places with contact data and reviews Turn any Google Maps search into structured records: business name, category, address, phone, website, rating, review count, opening hours, price range, coordinates, and up to 5 images per place. Optionally pull the top 10 reviews for each listing, billed at a flat $0.003 per place with no volume tiers. - Returns name, address, phone, website, rating, review count, hours, and coordinates per place. - Optional review scraping adds the top 10 reviews for each location. - Works through API, CLI, or MCP server, at a flat $0.003 per place. ### FAQ Q: How much does it cost to scrape Google Maps data? A: Pricing is a flat $0.003 per place with no volume tiers. A 1,000 place run costs $3.00, 10,000 places cost $30, and no API key or Google credentials are required. Q: How many results can you get from one Google Maps search? A: Google Maps itself returns roughly 120 places per search query, so a single run tops out around that number. To collect more, run narrower searches, for example one query per neighborhood or per business category. ## 30-Day Trend Intel Canonical: https://datapika.com/actors/30days-trend-intel Run at: https://apify.com/openclawai/30days-trend-intel Price: $0.35 per report Updated: 2026-08-29 Cross-platform trend reports with engagement scores and sentiment analysis Run one topic query across 13+ platforms, including Reddit, TikTok, Instagram, YouTube, Hacker News, GitHub, and Bluesky, over any window from 7 to 90 days. Each report returns ranked candidates with engagement scores, cross-source clusters, per-platform result counts, and an optional AI synthesis with executive summary, key findings, sentiment, and momentum, delivered as JSON or Markdown. - Searches 13+ platforms from a single topic input, with optional per-platform source filtering. - Four report tiers, from Quick Scan to Full Brief, free to run from August 31, 2026 with only Apify compute billed. - Optional AI synthesis adds executive summary, key findings, sentiment, and momentum. ### FAQ Q: How much does a cross-platform trend report cost? A: From August 31, 2026 the actor is free to run; you pay only for the Apify platform compute each run uses. Tiers range from Quick Scan to Full Brief, and you choose the depth per run, so a quick check and a deep briefing use the same tool. Q: What platforms does a 30 day trend analysis cover? A: Reports pull from Reddit, TikTok, Instagram, YouTube, Hacker News, GitHub, Threads, Pinterest, Polymarket, and Bluesky. You can search all of them at once or filter to specific sources, over a window of 7, 14, 30, or 90 days. ## Google Flights Scraper Canonical: https://datapika.com/actors/google-flights-scraper Run at: https://apify.com/openclawai/google-flights-scraper Price: $0.0015 per itinerary Updated: 2026-08-29 Scrape Google Flights itineraries with prices, airlines, durations, stops Runs route and date queries against Google Flights and returns each itinerary with total price, airline names and IATA codes, stop count, trip time in minutes, ISO departure and arrival times, and per-segment airports, times, and aircraft. Covers one-way, round-trip, and multi-city searches in four cabin classes, with results in 30+ languages and CO2 emissions reported for every itinerary. - Full itinerary detail: price, airlines, stops, duration, and every segment's airports and aircraft. - Handles one-way, round-trip, and multi-city trips across four cabin classes and passenger types. - Live Google Flights data at $0.0015 per itinerary, no API key required. ### FAQ Q: How much does it cost to scrape Google Flights data? A: This scraper charges $0.0015 per itinerary, which works out to $1.50 for 1,000 results. Google typically shows 30 to 80 itineraries per route and date query, so a single search usually costs between 5 and 12 cents. Q: What data does a Google Flights scraper return? A: Each itinerary includes the total price, airline names and IATA codes, stop count, total trip time in minutes, ISO departure and arrival times, and the airports, times, and aircraft for every segment. It also reports each itinerary's CO2 emissions in grams next to the typical figure for that route, with results available in 30+ languages. ## Hiring Signals Scanner Canonical: https://datapika.com/actors/indeed-ziprecruiter-scraper Run at: https://apify.com/openclawai/indeed-ziprecruiter-scraper Price: $0.0005 per record Updated: 2026-08-29 Track hiring signals and fresh jobs on Indeed and ZipRecruiter Feed in company names to get active posting counts on Indeed and ZipRecruiter, live ZipRecruiter job totals, and full profiles with industry, headcount, revenue, HQ, founding year, and rating. Keyword mode returns fresh job rows across both boards with structured salary data, direct application links, and exact UTC posting timestamps. - Live ZipRecruiter job totals plus per-board active posting counts for every company you check - Hour-precise posting timestamps with separate repost detection for true 24 and 48 hour filtering - Full company profiles: industry, headcount, revenue, HQ, founding year, rating, at $0.0005 per record ### FAQ Q: How can I check if a company is actively hiring on Indeed or ZipRecruiter? A: Feed company names into the scraper and each record returns a has_active_postings flag, employer-verified posting counts per board, and ZipRecruiter's live total of open jobs. At $0.0005 per record, checking 1,000 companies costs $0.50. Q: How do I filter scraped job postings to only the last 24 or 48 hours? A: Each job row carries an exact UTC posting timestamp where the board exposes one, with hour precision on ZipRecruiter, and company records include a posted_last_48h_count. Reposts are flagged separately in a reposted_at field, so a bumped older ad never counts as a genuinely new posting. ## Ads Transparency Scraper Canonical: https://datapika.com/actors/google-ads-transparency-scraper Run at: https://apify.com/openclawai/google-ads-transparency-scraper Price: $0.001 per result Updated: 2026-08-29 Extract advertiser creatives and reach data from Google Ads Transparency Look up any advertiser in the Google Ads Transparency Center by brand name, domain, or advertiser ID and get every creative as structured data: format, first and last shown dates, days active, image and video URLs, legal entity name, and per-country reach estimates. Coverage spans 230+ countries and every Google surface, including Search, YouTube, Display, Shopping, and Maps, at $0.001 per result. - Search by brand name, website domain, advertiser ID, or Transparency Center URL. - Returns text, image, and video creatives with run dates and days active. - Per-country reach estimates and surface breakdowns across Search, YouTube, Display, Shopping, and Maps. ### FAQ Q: How do I see all the ads a company is running on Google? A: Enter the brand name, website domain, or advertiser ID and the scraper pulls every creative from the Google Ads Transparency Center, up to 2,000 ads per advertiser. Each result includes the ad format, first and last shown dates, days active, preview URLs, and the advertiser's registered legal name. Q: How much does it cost to scrape the Google Ads Transparency Center? A: Results cost $0.001 each, which works out to $1.00 per 1,000 ads, with no separate fee to start a run. You can filter by ad format and region so you only pay for the results you need. # Guides ## How to scrape LinkedIn job postings without logging in Canonical: https://datapika.com/scrape/linkedin-jobs Updated: 2026-08-29 Scrape LinkedIn jobs by running Datapika's job-board-scraper on Apify: type a job title, set a location, pick linkedin as the source, and the run streams rows into a dataset within 5 to 20 seconds. No LinkedIn account, session cookie, or partner API key is involved, because the scraper reads the same public search pages a logged-out visitor sees. Each row carries title, company, location, job URL, seniority level, and, with linkedinFetchDescription enabled, the full posting text and direct apply link. You pay $0.005 per job delivered, and the actor has 27,458 runs from 2,471 users with a 5.0 rating. ### Can you scrape LinkedIn jobs without logging in or using cookies? Yes, and that is the whole design of the LinkedIn source in Datapika's job-board-scraper. LinkedIn publishes its job search results on public pages that a signed-out visitor can open, and the scraper requests those pages directly. It never asks for a LinkedIn username, password, li_at session cookie, or browser extension, so nothing you own is exposed to a restriction notice, and there is no cookie to refresh when it expires. The scraper's documentation states that it does not log in to any account, does not touch private candidate or recruiter data, and does not circumvent authentication; it only collects postings employers published to be found. That also means it cannot reach anything a session unlocks, such as submitting an Easy Apply form or reading applicant details. What you get is the public posting: title, company, location, date, seniority, and the description text. For most sourcing, lead generation, and market research pipelines that is exactly the set that matters. - No LinkedIn account, password, or li_at cookie in the input; the only required field is searchTerm - Reads public job search pages, the same ones a signed-out visitor loads - Nothing behind a login is touched: no Easy Apply submission, applicant lists, or member profiles - No browser extension and no session to refresh when LinkedIn rotates cookies - The default residential proxy carries the requests, so your office IP is never shown to LinkedIn ### What fields does a LinkedIn job scraper return? Every LinkedIn row lands in the dataset with the same flat schema the other seven boards use, plus one field the README ties to LinkedIn alone: job_level, the seniority label LinkedIn attaches to a posting. Alongside it you get the source job id, title, company, location, date_posted, job_type, is_remote, the site value linkedin, and, when the posting shows them, job_function, listing_type, company_industry, company_url, and company_logo. Salary fields have a salary_source marker: direct_data when LinkedIn shows a range, or description when the number was parsed from the text, and enforceAnnualSalary converts hourly and monthly figures to yearly. Two fields depend on linkedinFetchDescription: description, delivered in Markdown by default or HTML on request, and job_url_direct, the employer's own apply page when one exists. Emails found in the description are pulled into their own array. Each row also records search_term, matched_search_term when you run several queries, and a scraped_at ISO timestamp. The table below marks which fields LinkedIn actually fills. - job_level is LinkedIn-only: the seniority tag LinkedIn shows on the posting - Core row: id, title, company, location, job_url, date_posted, job_type, is_remote, site - Salary: salary_min, salary_max, salary_currency, salary_interval, with salary_source telling you whether it came from board data or parsed text - linkedinFetchDescription adds description (Markdown or HTML) and job_url_direct - emails array extracted from the description text - search_term, matched_search_term, and scraped_at on every row ### What are the quirks of scraping LinkedIn compared with other boards? LinkedIn behaves differently from Indeed in four ways that shape how you configure a run. First, LinkedIn rate-limits search traffic at roughly 100 results per IP address, which is why the proxy configuration defaults to Apify residential proxies; a single fixed IP hits that ceiling quickly on a large sweep, while rotating residential IPs spread the requests. Second, the search listing does not include the posting body, so linkedinFetchDescription triggers extra requests per job to collect the description and direct apply URL, making LinkedIn runs slower than the same query on Indeed. Leave it off when titles and companies are enough. Third, LinkedIn rejects the combination of hoursOld and easyApply in one search; pick one, or run two searches. Fourth, LinkedIn is the only board with company targeting, so linkedinCompanyIds can restrict results to named employers. Because LinkedIn is global, countryIndeed has no effect on it; steer by location and distance, which defaults to 50 miles. Rows still stream into the dataset 5 to 20 seconds after the run starts. - Rate limit: about 100 results per IP, so residential proxies are the default and the right choice for large runs - linkedinFetchDescription is off by default; enabling it adds requests and time but returns description and job_url_direct - hoursOld and easyApply cannot be combined on LinkedIn; Indeed has its own filter conflicts - linkedinCompanyIds restricts a search to specific employers, a LinkedIn-only feature - countryIndeed is ignored for LinkedIn; use location and distance (default 50 miles) - maxResults caps at 100 per board per search term, and searchTerms allows up to 5 queries per run ### How do you monitor specific companies' LinkedIn job postings? Recruiters, sales teams, and analysts often care about a handful of employers rather than a keyword. LinkedIn is the one board where the scraper can pin a search to named companies: pass the numeric IDs LinkedIn assigns to their company pages in linkedinCompanyIds, and the run ignores every other employer. The documented example is ["1441", "2382910"] with searchTerm product manager and location San Francisco. Pair that with hoursOld set to 24 or 168 and a daily Apify schedule, and each run returns only the roles those companies opened since the last one, which is the raw material for a competitor hiring tracker or a buying-signal feed. Use searchTerms with up to five titles to cover a whole function, since every row is tagged with matched_search_term. Because output is deduplicated across boards, adding Indeed to the same run brings in the company size, revenue, and rating columns shown in the README's sample Indeed row without doubling the row count. The full board list is at datapika.com/scrape. - linkedinCompanyIds accepts a list of numeric company IDs, for example ["1441", "2382910"] - hoursOld 24 on a daily schedule returns only postings opened since yesterday - searchTerms takes up to 5 titles and stamps matched_search_term on every row - Cross-board deduplication means a posting found on both LinkedIn and Indeed appears once in the dataset - Attach a webhook to push each finished run into Slack, Sheets, or a CRM ### How do developers and AI agents pull LinkedIn jobs from Datapika? Three paths lead to the same dataset. In the Apify Console you fill the form, run, and download JSON, CSV, Excel, XML, or RSS from the Output tab. From code, POST your input to the run endpoint for openclawai~job-board-scraper with your Apify token, or call run-sync-get-dataset-items to receive the job rows in a single response; the Python and Node clients wrap both in one call. For agents, add https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper to Claude Desktop, Cursor, or any MCP client, and the assistant can ask for remote senior React roles posted on LinkedIn in the last 24 hours and reason over typed rows instead of HTML. Set a maximum total charge on the run in Console and keep maxResults low while testing. Runs default to 4 GB of memory, which matters only when you add the browser-driven boards; a LinkedIn-only run finishes quickly. Zapier, Make, and n8n can trigger scheduled runs and route rows onward. - Console: form in, dataset out as JSON, CSV, Excel, XML, or RSS - REST: POST to api.apify.com/v2/acts/openclawai~job-board-scraper/runs, or run-sync-get-dataset-items for one call - Python: pip install apify-client, then client.actor("openclawai/job-board-scraper").call(run_input={...}) - MCP: mcp.apify.com with the job-board-scraper tool enabled for Claude, ChatGPT, or Cursor - Schedules plus webhooks turn one query into a daily LinkedIn feed ### Steps 1. Open apify.com/openclawai/job-board-scraper, enter a searchTerm such as "software engineer" and a location, and set sites to ["linkedin"]. 2. Turn on linkedinFetchDescription if you need the full posting and direct apply URL, add linkedinCompanyIds or hoursOld to narrow the sweep, and keep the default residential proxy. 3. Run it; LinkedIn rows start landing in the dataset within 5 to 20 seconds, and you can export JSON, CSV, or Excel or read them through the API. 4. For agents, connect https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper in your MCP client so the assistant can run LinkedIn searches on demand. ### FAQ Q: Does LinkedIn have an official API for job listings? A: Not for reading or searching postings. LinkedIn's developer documentation, updated June 2026, lists only three self-serve permissions: sign-in profile, email, and sharing posts. Every talent product, including Premium Job Posting, Apply Connect, and Recruiter System Connect, requires a partner application through LinkedIn Talent Solutions, and those integrations are built for posting jobs to LinkedIn or syncing applicants with an ATS, not for querying the job inventory. One August 2026 analysis puts Talent-track approval at 4 to 6 months minimum. Scraping the public search pages is the practical path for developers. Q: Do I need a LinkedIn login or cookies to scrape jobs with Datapika? A: No. The input has no field for a username, password, or session cookie, and the scraper never signs in. It reads the public job search pages that any signed-out visitor can open, through the default residential proxy, so your own LinkedIn account and IP address are never involved. The trade-off is scope: only what LinkedIn shows publicly is returned, which covers title, company, location, seniority, salary when listed, and the full description. Q: How much does it cost to scrape LinkedIn jobs? A: $0.005 per job delivered, so 1,000 LinkedIn postings cost $5, billed only for rows that reach your dataset; a search that returns nothing costs nothing beyond Apify's small run start fee. The ceiling for a run is maxResults times the number of boards times the number of search terms, so a LinkedIn-only sweep with maxResults 100 and 3 terms tops out at 300 rows. LinkedIn's per-IP limit and filter matches usually bring the actual count lower. Q: Why does a LinkedIn run return fewer than 100 results? A: Three causes cover most cases. LinkedIn throttles at roughly 100 results per IP, so a single fixed IP can hit the ceiling early; keep the residential default. Filter conflicts are the second: LinkedIn will not combine hoursOld with easyApply, so drop one. Third, the query simply has fewer matches within the distance radius, which defaults to 50 miles. Widen location, add searchTerms, or paginate with offset to reach more postings. Q: Is it legal to scrape LinkedIn job postings? A: The scraper collects only public postings that employers published to be found; it does not log in, bypass authentication, or read member profiles or applicant data. Job listings are business information, though a description may include a recruiter's email, which lands in the emails array and falls under privacy law such as GDPR. You are responsible for using the output in line with LinkedIn's terms and the laws that apply to you, and for handling any personal contact details accordingly. ## Scrape Indeed job postings with salary and company data Canonical: https://datapika.com/scrape/indeed-jobs Updated: 2026-08-29 Indeed is the board to start with when you need job listings in bulk. Datapika's job board scraper on Apify reads Indeed's public search pages with no API key or login, returns each posting as a structured row with salary range, company size, revenue, rating and full description, and streams the first Indeed rows into your dataset 5 to 20 seconds after the run starts. Set countryIndeed to target any supported Indeed country site. You pay $0.005 for each job delivered, which puts 1,000 Indeed listings at $5. The scraper has 27,458 runs, 2,471 users and a 5.0 rating. ### What fields does an Indeed job scraper return? Every Indeed posting becomes one dataset row, and the sample row in the actor's documentation is an Indeed posting, which shows how much of the schema Indeed fills in. The salary block carries salary_min, salary_max, salary_currency and salary_interval, plus a salary_source flag that tells you whether the range came from Indeed's own structured pay data or was parsed out of the description text; set enforceAnnualSalary to convert hourly and monthly figures to yearly equivalents before you compare roles. The company block is where Indeed stands out: employee count with a readable label such as 1001-5000, revenue with a label such as $1B+, an employer rating with its review count, industry, country, logo URL and a short description. In that sample, an Indeed row for a Senior Software Engineer at JPMorganChase shows a 120,000 to 185,000 USD yearly range, a 3.9 rating from 21,432 reviews and a 10,000+ headcount. The full description arrives by default in Markdown, or HTML if you set descriptionFormat, and any email addresses found in it are extracted into a separate emails field. - salary_min, salary_max, salary_currency, salary_interval and salary_source (direct_data or description) - company_num_employees and company_revenue, each with a human-readable label field - company_rating, company_reviews_count, company_industry, company_country, company_logo and company_description - job_url to the Indeed listing plus job_url_direct to the employer's own careers page when Indeed exposes it - description in Markdown or HTML without a separate fetch switch (LinkedIn needs linkedinFetchDescription for the same field) - emails extracted from the description text, useful for recruiter contact ### How do you set the Indeed country for a search? Indeed runs a separate site for each country, and the scraper picks which one to query through the countryIndeed input. It defaults to usa; the documented codes include uk, canada, australia, germany, france, india, singapore and uae, and most other Indeed country sites are accepted as well. The location field then narrows results inside that country, so a run with countryIndeed set to uk and location set to Manchester returns Manchester postings from Indeed's UK site. Leaving countryIndeed on the default while typing a London location is the most common configuration mistake: the scraper dutifully asks the US site for London jobs, so check the country before you blame the query. The same value also drives Glassdoor's country targeting, so a UK run that includes both boards needs only one setting. Pair the country with distance, which defaults to a 50-mile radius, to control how far outside the named city the search reaches, and turn on enforceAnnualSalary when you plan to compare pay across countries with different salary conventions. - countryIndeed defaults to usa and accepts codes such as uk, canada, australia, germany, france, india, singapore and uae - location narrows inside the chosen country: a city, a region or the word Remote - distance sets the search radius in miles, default 50 - One countryIndeed value targets both Indeed and Glassdoor in the same run - salary_currency accompanies each salary range, so mixed-country datasets stay comparable - Pick the country first, then the location; the country chooses the site and the location narrows within it ### What are Indeed's quirks compared with the other boards? Indeed is the only board in the actor's site notes flagged as having no rate limiting, which is why those notes call it the best choice for large scrapes. It also skips the stealth browser, so its rows land in the dataset 5 to 20 seconds after the run starts, while Glassdoor, ZipRecruiter, Bayt and Naukri are fetched through a real browser and typically take 1 to 3 minutes. Three limits are worth knowing. First, Indeed cannot combine hoursOld with jobType, isRemote or easyApply in one search; if you set both, the run returns fewer rows than expected, so filter by recency in one run and by role type in another, or post-filter the dataset. Second, date_posted is a calendar date with no time of day, so a 24-hour window is day-precise rather than hour-precise. Third, job_level is a LinkedIn field and comes back null on Indeed rows, so use title keywords for seniority instead. A listing_type field flags sponsored placements, which lets you separate paid promotion from organic postings when you count who is hiring. - No rate limiting on Indeed, so large maxResults values and repeated scheduled runs are fine - No stealth browser step: first Indeed rows in 5 to 20 seconds, not minutes - hoursOld cannot be combined with jobType, isRemote or easyApply on Indeed - date_posted is day-precise (YYYY-MM-DD), not an hour-level timestamp - job_level is null on Indeed rows; skills and experience_range are Naukri-only - listing_type marks sponsored listings ### How do you scrape Indeed jobs at volume? A single run returns up to maxResults rows per board per search term, with maxResults capped at 100 and searchTerms capped at 5, so an Indeed-only run tops out at 500 rows before deduplication. To go beyond that, page with the offset input, which skips the first N results, or split the job across locations, job types and countries and let Apify schedules run the pieces. Each row carries matched_search_term, so a multi-term OR search such as data engineer, analytics engineer and ETL developer stays traceable after the results are merged. For a daily feed, keep hoursOld at 24 and run once a day; because you pay only for rows delivered, a morning that turns up nothing costs no result events. Runs default to 4 GB of memory, which matters mainly when browser-based boards are included; an Indeed-only run is light. Cap each run with a maximum total charge in the Apify Console as a hard budget, and keep maxResults at the default of 20 while you tune the query before raising it to 100. - maxResults up to 100 per term and searchTerms up to 5: 500 Indeed rows per run before dedup - offset skips the first N results for pagination across repeated runs - matched_search_term tags every row with the query that found it - hoursOld 24 plus a daily schedule gives a fresh-postings feed - Zero-result runs produce no result events, only Apify's small start fee - A maximum total charge on the run caps spend before it starts ### How do you call the Indeed scraper from code or an AI agent? The actor has three entry points. In the Apify Console you fill the form, run it and download the dataset as JSON, CSV, Excel, XML, HTML or RSS from the Output tab. From code, POST your input to the runs endpoint for openclawai~job-board-scraper, or call run-sync-get-dataset-items to get the rows back in one synchronous request; the Python and JavaScript clients wrap the same calls behind client.actor("openclawai/job-board-scraper").call(). For agents, register the scraper as an MCP tool via mcp.apify.com, and a Claude, ChatGPT or Cursor session can ask for remote data analyst jobs on Indeed posted in the last day and reason over the structured rows directly. Zapier, Make and n8n can run the scraper on a schedule and push new rows to Google Sheets, Airtable, Slack or a CRM, and a webhook can kick off downstream processing when a run finishes. When one board is not enough, the same input can sweep LinkedIn, Glassdoor, ZipRecruiter, Naukri and Bayt in the same run with cross-board deduplication; the full board list is at datapika.com/scrape. - Console export: JSON, CSV, Excel, XML, HTML table or RSS - REST: POST to /v2/acts/openclawai~job-board-scraper/runs, or run-sync-get-dataset-items for a single call - Python and Node clients: pip install apify-client or npm i apify-client - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper - Zapier, Make and n8n schedules, plus webhooks on run completion ### Steps 1. Open apify.com/openclawai/job-board-scraper, enter a searchTerm such as data analyst (or up to 5 searchTerms) and a location. 2. Set sites to indeed only, set countryIndeed to the country whose Indeed site you want (usa, uk, canada, india and so on) and choose maxResults up to 100. 3. Add filters: either hoursOld for recency, or jobType, isRemote and easyApply for role type, but not both together on Indeed. Turn on enforceAnnualSalary if you will compare pay. 4. Run it; Indeed rows appear within 5 to 20 seconds. Export from the Output tab or read the dataset through the API or MCP. ### FAQ Q: Does Indeed have an official API? A: Not for reading job postings. Indeed's Publisher Job Search API was shut down, its old developer.indeed.com documentation now redirects to partners.indeed.com, and the APIs Indeed documents today, such as the GraphQL Job Sync API, exist for approved ATS partners to push their own postings into Indeed rather than to search listings. There is no self-serve key to apply for. Datapika reads the public search pages instead and returns the same fields you see in a browser. Q: What does it cost to pull 1,000 Indeed jobs? A: One thousand Indeed rows cost $5, billed per delivered job at $0.005 on a pay-per-event basis. Rows removed by deduplication never reach your dataset, so they are not charged, and a run that returns nothing costs only Apify's small start fee. Apify's free plan includes $5 of monthly credit, roughly enough for 1,000 jobs, and a maximum total charge set on the run caps spend before it starts. Q: Why does Indeed return fewer jobs than maxResults? A: Usually one of three things. Indeed cannot combine hoursOld with jobType, isRemote or easyApply in the same search, so that combination silently narrows results; run recency and role-type filters separately. Duplicates are removed when the same posting appears on several boards in a multi-board run. Or the country and location disagree, such as a London location on the default usa site. Widen the radius or drop hoursOld to check. Q: Can I scrape Indeed jobs from the UK, Canada, India or other countries? A: Yes. Set countryIndeed to the country code for the Indeed site you want; usa is the default, uk, canada, australia, germany, france, india, singapore and uae are documented codes, and most other Indeed country sites are accepted too. Then put a city or region in location. For India-specific fields such as skills and experience_range, add Naukri to the same run alongside Indeed. Q: How fresh are Indeed postings, and can I watch when a specific company starts hiring? A: Indeed exposes a posting date, not a time, so hoursOld set to 24 gives you a day-precise window; run it daily to keep a rolling feed of new listings. For a company-by-company view, the hiring signals use case uses the companion Indeed and ZipRecruiter scraper, which takes a list of employers and reports which ones posted in the last 48 hours, with the titles and pay ranges attached. ## How to scrape Glassdoor job listings and their salary ranges Canonical: https://datapika.com/scrape/glassdoor-jobs Updated: 2026-08-29 Datapika's job board scraper reads Glassdoor's public search results and returns each listing as a JSON row carrying the pay band Glassdoor attaches to it (low, high, currency, pay period), the posting date, the employer's Glassdoor overview link and the full description. You pick a country code such as usa or uk, add up to 5 search terms and cap results at 100 per term. Glassdoor is fetched through a real browser, so rows land within 1 to 3 minutes. Pricing is $0.005 per delivered job; the actor shows 27,458 runs and a 5.0 rating (August 29, 2026). ### What Glassdoor fields does the scraper return? A Glassdoor row is identified by an id prefixed gd- and a site value of glassdoor, with title, company, location, job_url, date_posted and an is_remote flag. The compensation block is the reason to pick this board: salary_min, salary_max, salary_currency and salary_interval mirror the pay band Glassdoor attaches to the listing header, and salary_source reads direct_data when that band exists. company_url points at the employer's overview page on the Glassdoor country site, company_logo holds the square logo, and listing_type records the sponsorship level, so paid placements can be separated from organic postings. The description is requested for each job and delivered as Markdown by default or HTML on request, and any email addresses found inside it are copied into emails. Fields that Glassdoor's search results do not expose stay null on these rows: company_rating, company_reviews_count, company_num_employees, company_revenue, company_industry, job_type and job_url_direct. If you want a rating next to the pay band, keep Indeed in the same run, because Indeed rows carry company_rating and company_reviews_count. - Identity: id (gd- prefix), site, title, company, location, job_url, date_posted, is_remote - Pay: salary_min, salary_max, salary_currency, salary_interval, salary_source - Employer: company_url (Glassdoor overview page), company_logo, listing_type - Content: description in Markdown or HTML, emails parsed from the text - Null on Glassdoor rows: company_rating, company_reviews_count, company size and revenue, job_type, job_url_direct - Every row also carries search_term, matched_search_term and a scraped_at UTC timestamp ### How do Glassdoor salary bands work in the output? Glassdoor attaches a pay band to most listings, whether the employer supplied it or Glassdoor estimated it, and the scraper records the low and high points of that band as salary_min and salary_max. salary_interval tells you whether the figures are yearly, monthly or hourly, which matters for retail and healthcare roles that Glassdoor quotes per hour. Turn on enforceAnnualSalary and hourly bands are multiplied by 2080 and monthly bands by 12 before they reach the dataset, so a $28 to $35 hourly nursing job and a $95,000 to $120,000 salaried one sort on one axis. salary_currency is the currency code Glassdoor reports for the band, so uk and germany runs return GBP and EUR figures rather than converted dollars. When a listing carries no band and countryIndeed is usa, the scraper scans the description for dollar amounts and sets salary_source to description; outside the US that fallback is not attempted, so a missing salary_min on a uk run means Glassdoor showed no pay for that listing. - salary_min and salary_max are the low and high points of the band Glassdoor attaches to the listing - salary_interval is yearly, monthly or hourly; enforceAnnualSalary converts hourly (x 2080) and monthly (x 12) to yearly - salary_currency is the code Glassdoor reports for the band, so a uk run returns GBP figures - salary_source is direct_data for Glassdoor bands, description for the US text fallback, null when there is no pay - Missing pay on a non-US row means Glassdoor displayed no band for that listing - README example: searchTerm machine learning engineer, location London, countryIndeed uk, enforceAnnualSalary true ### Which country, location and filter settings does Glassdoor accept? Glassdoor runs one site per country, so the countryIndeed input decides which one is queried: usa is the default, and uk, canada, australia, germany, france, india, singapore and uae are among the accepted codes, along with most common country names and 2-letter ISO codes. The location string is resolved against Glassdoor's own place index, which returns a city, state or country entity, so write it the way Glassdoor does, for example Nashville, TN or London. The distance input is not applied to this board because Glassdoor searches by place entity rather than radius. Set isRemote true and the scraper skips the place lookup and queries Glassdoor's remote location entity directly. hoursOld is honored, but Glassdoor filters by whole days, so 24 becomes 1 day and 168 becomes 7 days; values under 24 still count as 1 day. jobType maps to Glassdoor's employment type filter and easyApply restricts results to postings that accept applications on Glassdoor. Results arrive newest first, 30 per page, up to the maxResults cap of 100 per search term, and offset skips whole pages of 30. - countryIndeed picks the Glassdoor country site; examples: usa, uk, canada, australia, germany, france, india, singapore, uae - location resolves to a Glassdoor city, state or country entity; distance is not applied - isRemote true searches Glassdoor's remote location entity without a place lookup - hoursOld rounds down to whole days with a 1-day minimum (24 = 1 day, 168 = 7 days) - jobType (fulltime, parttime, contract, internship, temporary) and easyApply both work on Glassdoor - Newest first, 30 per page, maxResults up to 100 per term, offset skips whole pages ### What Glassdoor quirks should you plan for? Glassdoor's anti-bot wall rejects plain HTTP clients, so every request for this board is issued from a real browser session. That is why Indeed and LinkedIn rows appear 5 to 20 seconds after start while Glassdoor rows take 1 to 3 minutes, and why runs default to 4 GB of memory. Each description is a separate lookup; when a search term's time budget is nearly spent the scraper skips the remaining description fetches so rows still land with their pay bands intact. Cross-board deduplication keeps the first row to arrive per title and company, and since Indeed usually finishes first, a job on both boards is delivered as the Indeed row. Run Glassdoor alone when you need its pay band on every listing, or with Indeed when you want ratings. date_posted is a calendar date derived from Glassdoor's days-old counter, with no time of day. Glassdoor's site migration in April 2026 briefly made every search return zero rows; v1.0.17 on 2026-04-20 restored results, with a Nashville, TN software engineer search returning 10 jobs from Deloitte, KPMG, PwC and Amazon. - Browser-based fetch: Glassdoor rows land 1 to 3 minutes after start, Indeed and LinkedIn in 5 to 20 seconds - Descriptions are fetched per job and skipped near the time budget, keeping salary data - Dedup keeps the first arrival per title and company, so Indeed rows usually win over Glassdoor duplicates - Run sites: ["glassdoor"] alone to get a Glassdoor pay band on every row that has one - date_posted has day precision only - The April 2026 Glassdoor site migration was handled in v1.0.17; results resumed on 2026-04-20 ### How do you run the Glassdoor scraper from the console, the API or an AI agent? No Glassdoor account, partner agreement or browser extension is involved. In the Apify Console you enter a search term, set sites to glassdoor, choose the country, and read the dataset in the browser or download it as JSON, CSV, Excel, XML or RSS. From code, POST the same input to https://api.apify.com/v2/acts/openclawai~job-board-scraper/runs with your Apify token, or use run-sync-get-dataset-items for one blocking call that returns the rows; the Python and Node clients wrap both. Agents get the scraper as a tool through the Apify MCP server at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper, so a Claude or Cursor session can ask for Glassdoor pay bands for a title in a city and reason over the rows. Billing is $0.005 per job delivered: a 20-row test costs $0.10 and 1,000 rows cost $5.00, and an empty result carries no per-job charge. Apify's free plan includes $5 of monthly credit, roughly 1,000 Glassdoor jobs. Zapier, Make and n8n can schedule the run and push new rows to Sheets, Slack or a CRM. The same input can sweep the 7 other boards alongside Glassdoor. - Console: searchTerm, sites ["glassdoor"], countryIndeed, then export JSON, CSV, Excel, XML or RSS - REST: POST to /v2/acts/openclawai~job-board-scraper/runs, or run-sync-get-dataset-items for one call - Python: pip install apify-client, then client.actor("openclawai/job-board-scraper").call(run_input=...) - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper - Billing only for delivered rows; an empty result carries no per-job charge - Zapier, Make and n8n schedules push fresh Glassdoor rows into Sheets, Slack or HubSpot ### Steps 1. Open https://apify.com/openclawai/job-board-scraper, type a search term such as data analyst, set sites to glassdoor, and set countryIndeed to the Glassdoor country you want (usa, uk, canada, germany and others). 2. Add a location written the way Glassdoor shows it (London or Nashville, TN), or set isRemote true, then pick maxResults up to 100, an optional hoursOld window in whole days, jobType or easyApply, and enable enforceAnnualSalary if you want hourly bands converted to yearly. 3. Start the run and wait 1 to 3 minutes for the browser-based fetch; rows stream into the dataset with salary_min, salary_max, salary_currency, salary_interval and the description once the board finishes. 4. Export JSON, CSV or Excel from the dataset, call the same input through the Apify API or Python client for a pipeline, or add the MCP URL so an agent can query Glassdoor pay bands on demand. ### FAQ Q: Does Glassdoor have an official API? A: Not one you can sign up for. Glassdoor retired its public developer API and issues no new developer keys in 2026; the remaining API page on help.glassdoor.com describes Customer Insights, a B2B subscription product for HR teams, not a job-data feed (DEV Community, April 29, 2026; JobsPipe, updated August 29, 2026). Datapika reads the public search pages instead, so there is no key, waitlist or partner contract to obtain. Q: Do Glassdoor rows include the company rating and review count? A: No. company_rating and company_reviews_count are null on rows where site is glassdoor, because the search results the scraper reads do not carry them. You do get company_url, which links to the employer's Glassdoor overview page, and company_logo. Ratings are populated on Indeed and Naukri rows, so a combined Glassdoor plus Indeed run gives you pay bands from one board and ratings from the other. Q: Why does a Glassdoor run return fewer jobs than maxResults, or none at all? A: Usual causes are a location string Glassdoor cannot resolve, a country code that does not match the location, or a narrow hoursOld window; remember hoursOld works in whole days on this board. Duplicates found on faster boards are also removed before Glassdoor's rows arrive. Run Glassdoor alone, loosen the filters, and check that countryIndeed matches the city you typed. An empty result carries no per-job charge. Q: Which countries can I scrape Glassdoor jobs from? A: Any country with its own Glassdoor site. The countryIndeed input accepts codes such as usa, uk, canada, australia, germany, france, india, singapore and uae, plus most common country names and 2-letter ISO codes. The salary_currency on each row is the code Glassdoor reports for that listing, so bands are not converted to dollars. One run queries one country; schedule separate runs per market if you need several. Q: Can an AI agent or a scheduled pipeline pull Glassdoor salary data? A: Yes. Add https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper to Claude Desktop, Cursor or any MCP client and the agent can run a Glassdoor search and read the salary_min, salary_max and salary_interval fields directly. For pipelines, schedule the actor in Apify Console with hoursOld set to 24 or 168, attach a webhook, and push new rows to Sheets, Airtable or a CRM through Zapier, Make or n8n. Q: Is scraping Glassdoor job listings allowed? A: The scraper only reads job postings that Glassdoor shows to any anonymous visitor. It does not log in, does not touch reviews, salaries submitted by employees, or any private profile data, and does not bypass authentication. You remain responsible for using the rows in line with Glassdoor's terms and with laws that apply to you, including GDPR when a description contains a recruiter's email address. ## How to scrape Google Jobs listings from their source boards Canonical: https://datapika.com/scrape/google-jobs Updated: 2026-08-29 Google Jobs is an aggregator, not a job board: it indexes postings that employers and boards like Indeed, LinkedIn and Glassdoor publish with JobPosting structured data, and shows a trimmed card for each. The same postings come from those source boards with fuller fields, so the way to scrape Google Jobs is to sweep the boards directly. Datapika's job board scraper on Apify does that in one run at $0.005 per job, returning salary, company rating, employee count and the full description per row. It has 27,458 runs and 2,471 users, with 381 active in the last 30 days. ### Why is Google Jobs an aggregator rather than a job board? Google Jobs does not host postings. Employers and job sites mark up their own pages with JobPosting structured data, ping Google through the Indexing API so Googlebot crawls them sooner, and Google folds the results into the jobs panel in Search. Google's own documentation recommends the Indexing API over sitemaps for job posting URLs for that reason. Every card in that panel therefore points back to an origin: an Indeed listing, a LinkedIn listing, a Glassdoor listing, or a company careers page. The card typically shows title, company, location and a posting age, sometimes a pay figure, and then hands you off to the source to read the rest. That is why scraping Google Jobs is a detour. The source board holds the full description, the direct apply link, company size, revenue and rating, and a salary field flagged as board data rather than guessed from text. Datapika's scraper reads those boards directly and merges duplicates into one row, which is the coverage Google Jobs advertises with the fields it hides. - Postings enter Google Jobs through JobPosting structured data on the employer's or board's own page (Google Search Central docs) - Google recommends the Indexing API over sitemaps for job URLs so new postings are crawled sooner - Each Google Jobs card links out to Indeed, LinkedIn, Glassdoor, ZipRecruiter or a careers page for the full posting - Google documents no read API for its jobs panel; Cloud Talent Solution is a separate product for building search over job content you upload - Datapika returns the source rows with salary_source, company_rating, company_num_employees and job_url_direct - Duplicates across boards are merged, so a job Google would show once still lands as one row ### How do you set up a Google Jobs style sweep in Datapika? Start with searchTerm, which is required, or searchTerms with up to 5 queries that run in sequence and merge, each row tagged with matched_search_term. Add a location such as New York, London or Remote; distance defaults to 50 miles around it. Choose sites. The default list is linkedin, indeed, glassdoor, zip_recruiter, bayt and naukri, which are the six boards currently returning results, and google stays selectable if you want to include it. If you do include it, put the exact phrase from the Google Jobs search box into googleSearchTerm; that input affects only the Google board. Set countryIndeed for Indeed and Glassdoor when the market is outside the US, for example uk, canada, australia, germany, france, india, singapore or uae. Then narrow with hoursOld (24 for a daily pull, 168 for weekly), isRemote, jobType (fulltime, parttime, contract, internship, temporary) and easyApply. maxResults defaults to 20 per board per term and caps at 100, so total rows are bounded by maxResults times boards times terms before deduplication. - searchTerm is required; searchTerms accepts up to 5 queries merged as an OR search - sites default: linkedin, indeed, glassdoor, zip_recruiter, bayt, naukri; add google explicitly - googleSearchTerm: paste the phrase from the Google Jobs UI, applies to the Google board only - countryIndeed steers Indeed and Glassdoor; LinkedIn is global - hoursOld, isRemote, jobType, easyApply, distance and offset narrow or page the results - maxResults 1 to 100 per board per term, default 20 ### What Google Jobs quirks should you know before running it? Google is the odd board in the list, and four behaviors matter. First, Google Jobs has no country switch of its own in this scraper; countryIndeed steers Indeed and Glassdoor, while Google follows the location string and the phrase you pass in googleSearchTerm. Second, that phrase should be copied verbatim from the Google Jobs search box, which the scraper's notes list as the query format that works best for that board. Third, since Google aggregates the other boards, any row it returns is likely to collide with an Indeed, LinkedIn or Glassdoor row for the same posting; the scraper deduplicates, so the merged row carries the richer source fields and Google adds little on its own. Fourth, google is off the default sites list since v1.0.49 (2026-08-03), so a default run never waits on it; the FAQ below covers why and what to run instead. Google accepts the shared filters isRemote, jobType and hoursOld, as the scraper's remote full-time last-24-hours example, which lists google among its sites, shows. - No country input for Google; countryIndeed applies to Indeed and Glassdoor only - googleSearchTerm should be the exact phrase from the Google Jobs search box - Rows from Google usually duplicate a source board row and are merged by the deduplicator - Removed from the default sites list in v1.0.49 (2026-08-03), still selectable - Shared filters isRemote, jobType and hoursOld apply to the Google board - LinkedIn, one of Google's sources, rate-limits around 100 results per IP, so residential proxy is the default ### Which boards reproduce Google Jobs coverage for a given market? Google Jobs is only as wide as its feeders, so pick the ones that match your market; the full board list lives at datapika.com/scrape. For the United States, run indeed, linkedin, glassdoor and zip_recruiter together; Indeed is the board the scraper's notes rate most reliable with no rate limiting, so it is the anchor for large pulls. For the UK, Canada, Australia, Germany, France, Singapore or the UAE, keep linkedin and set countryIndeed so Indeed and Glassdoor query the local site. For India, add naukri, which returns skills, experience_range, vacancy_count and work_from_home_type that no Google card shows. For the Gulf, add bayt, which supports searchTerm only. Speed differs by board: Indeed and LinkedIn rows land 5 to 20 seconds after start, while Glassdoor, ZipRecruiter, Bayt and Naukri go through a real browser and take 1 to 3 minutes. One run tops out at 100 results per board per term across 5 terms, or 4,000 rows over 8 boards, with offset for paging beyond that. - US: indeed, linkedin, glassdoor, zip_recruiter; ZipRecruiter covers the US and Canada only - UK, Canada, Australia, Germany, France, India, Singapore, UAE: set countryIndeed for Indeed and Glassdoor - India: naukri adds skills, experience_range, vacancy_count and work_from_home_type - Middle East: bayt, searchTerm filter only - Indeed and LinkedIn rows arrive in 5 to 20 seconds; browser-based boards take 1 to 3 minutes - Ceiling per run: 100 per board per term, 5 terms, 4,000 rows; page with offset ### How do developers and AI agents call the scraper instead of a Google Jobs API? There is no endpoint to query Google's jobs panel, so the scraper stands in for one. From the Apify Console you fill the form, run it, and download the dataset as JSON, CSV, Excel, XML or RSS. From code, POST the input to api.apify.com/v2/acts/openclawai~job-board-scraper/runs with your token and read the default dataset, or call run-sync-get-dataset-items to get the rows back in a single request. Python users install apify-client and call client.actor('openclawai/job-board-scraper').call(run_input=...), then list_items on the returned dataset id; Node users do the same with npm i apify-client. For agents, add https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper to Claude, ChatGPT, Cursor or any MCP client and the assistant can run a search and reason over the rows. Zapier, Make and n8n connect through the Apify app for scheduled pushes to Sheets, Airtable, Slack or a CRM, and a webhook on run finish triggers downstream processing. The listing shows 27,458 runs and a 5.0 rating from 3 reviews. - Console: form in, dataset out as JSON, CSV, Excel, XML or RSS - REST: POST /v2/acts/openclawai~job-board-scraper/runs, or run-sync-get-dataset-items for one call - Python and Node: apify-client with actor id openclawai/job-board-scraper - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper - Zapier, Make, n8n and webhooks for schedules and downstream pushes ### Steps 1. Open https://apify.com/openclawai/job-board-scraper, enter a searchTerm (or up to 5 searchTerms) and a location, and leave sites on the six-board default; add google explicitly only if you want to test it, with the Google Jobs phrase in googleSearchTerm. 2. Set countryIndeed for a non-US market, then add hoursOld 24 for a daily feed and isRemote or jobType as needed; keep maxResults at 20 for a first pass and raise it to 100 once the rows look right. 3. Run it. Indeed and LinkedIn rows land within 5 to 20 seconds; the browser-based boards follow in 1 to 3 minutes, deduplicated against each other. 4. Export JSON, CSV or Excel, read the dataset through the API, or attach the MCP URL to your agent and schedule the same input daily. ### FAQ Q: Does Google Jobs have an official API? A: No. Google offers no API that reads the jobs panel in Search. Its Search Central documentation covers only the inbound side: employers add JobPosting structured data and call the Indexing API so Googlebot picks up postings sooner. Cloud Talent Solution is a separate Google Cloud product that builds search over job content you upload to it, not a feed of what Google shows searchers. That leaves reading the source boards directly, which is what this page describes and what the scraper does. Q: Why does the google board currently return no results? A: Because of platform-side changes on Google's end that stopped the board returning rows. The v1.0.49 release on 2026-08-03 removed google from the default sites list so default runs no longer wait on it, and the site notes say it will resume when the platform stabilizes. It stays selectable in the meantime. Since Google only indexes postings that already live on Indeed, LinkedIn, Glassdoor and the other boards, sweeping those sources returns the same jobs with more fields, so nothing is lost while Google is dark. Q: Do I pay twice when Google Jobs and Indeed return the same posting? A: No. The deduplicator merges a job seen on several boards into a single row, and billing counts only rows delivered to the dataset, so a run costing $5 per 1,000 jobs never charges for the Indeed copy and the Google copy separately. A run that finds nothing costs nothing beyond Apify's small start fee. To keep repeat runs cheap, set hoursOld so each pull only fetches postings newer than the last one, and cap spend with a maximum total charge in Console. Q: Does googleSearchTerm replace searchTerm? A: No. searchTerm (or the searchTerms array) is required and drives every board. googleSearchTerm is an optional override applied to the Google board only, and the documentation recommends copying it straight from the Google Jobs search box because that phrasing produces the best results there. Because it is scoped to Google, it has no effect on the six default boards, so it can sit in a saved input while google stays deselected and come into play only when you add google back to sites. Q: Which boards should I select to replace a Google Jobs sweep for my country? A: Match the feeders to the market. LinkedIn is global. Indeed and Glassdoor follow countryIndeed, which accepts usa, uk, canada, australia, germany, france, india, singapore, uae and most other countries. ZipRecruiter covers the US and Canada, Naukri covers India and Bayt the Middle East. A US sweep is indeed, linkedin, glassdoor and zip_recruiter; a UK sweep is linkedin plus indeed and glassdoor with countryIndeed set to uk; an India sweep adds naukri for its skills and experience fields. Q: Can an AI agent use this scraper as a Google Jobs API? A: Yes. Add https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper to Claude, ChatGPT, Cursor or another MCP client and the agent gets a tool that takes searchTerm, location, sites and filters and returns typed rows. A prompt like 'find remote senior React jobs posted in the last 24 hours and list companies with salary ranges' maps to hoursOld 24 and isRemote true across linkedin and indeed. The same input works over the Apify REST API for scheduled runs. ## Scrape ZipRecruiter job postings without an API key Canonical: https://datapika.com/scrape/ziprecruiter-jobs Updated: 2026-08-29 ZipRecruiter has no public job search API: its ZipSearch feed was switched off on March 31, 2025, so the practical route is Datapika's job board scraper on Apify. Set sites to zip_recruiter, add a search term and a US or Canadian location, and every posting comes back as one JSON row with title, company, salary range, direct apply URL, and full description, no login or key required. Rows are billed at $0.005 each ($5 per 1,000), and the actor has logged 27,458 runs for 2,471 users. Export JSON, CSV, or Excel, or read the dataset over the API or MCP. ### What fields does a ZipRecruiter job row contain? Each ZipRecruiter posting lands in the dataset as a single item that shares its schema with every other board the actor supports, so a ZipRecruiter row and an Indeed row sit in the same CSV without column remapping. The core of the row is the job identity: a board-issued id, title, company, location, the job_url on ZipRecruiter, and job_url_direct pointing at the employer's own application page when the listing exposes one. Pay data is split into salary_min, salary_max, salary_currency, and salary_interval, with salary_source telling you whether the numbers came from the board's structured data or were parsed out of the description text; enforceAnnualSalary converts hourly and monthly figures to yearly so a $28 per hour warehouse role and a $58,000 salaried role compare on one axis. The description arrives as Markdown by default or HTML if you set descriptionFormat, and any email addresses found inside it are lifted into an emails array. Every row also records search_term, matched_search_term, and a scraped_at ISO timestamp. - Identity: id, title, company, location, site set to zip_recruiter, date_posted - Links: job_url on ZipRecruiter plus job_url_direct to the employer's careers page when available - Pay: salary_min, salary_max, salary_currency, salary_interval, and salary_source (direct_data or description) - Work terms: job_type (fulltime, parttime, contract, internship, temporary) and is_remote - Text: full description in Markdown or HTML, with extracted emails as a separate array - Provenance: search_term, matched_search_term, and scraped_at on every row ### What is different about scraping ZipRecruiter compared with other boards? ZipRecruiter behaves unlike the request-based boards in three ways you should plan around. First, coverage is limited to the United States and Canada, with no country switch equivalent to countryIndeed. Second, it is fetched through a real browser with an anti-bot warm-up phase, so ZipRecruiter rows typically land 1 to 3 minutes after the run starts, while Indeed and LinkedIn rows appear in 5 to 20 seconds. Results stream per board, so you can read the fast boards while ZipRecruiter is still loading, and the v1.0.49 release halved the warm-up retry budget to shorten unlucky runs. Third, ZipRecruiter rate-limits datacenter IPs, so keep proxyConfiguration on its residential default; runs also get 4 GB of memory for the browser session. On freshness, ZipRecruiter loads new listings once a day in a batch, so the youngest posting you see is often 12 to 48 hours old at scrape time. A 24-hour hoursOld window on ZipRecruiter alone will look thin, and 48 hours is the honest minimum. Unlike LinkedIn and Indeed, ZipRecruiter has no documented filter combinations that cancel each other out. - US and Canada only; there is no country parameter for ZipRecruiter - Browser-fetched with anti-bot warm-up: expect 1 to 3 minutes before rows appear - Datacenter IPs get rate-limited, so keep the residential proxy default - Daily batch loading means the youngest posting is often 12 to 48 hours old; use hoursOld of 48 or more - No documented restrictions on pairing hoursOld with isRemote, jobType, or easyApply, unlike LinkedIn and Indeed - Default memory is 4 GB to give the browser session headroom ### How do you get hour-precise posting times and repost flags from ZipRecruiter? The job board scraper records date_posted as a calendar date, such as 2026-07-07, with no time of day. ZipRecruiter's job pages carry more than that: an exact posting time and, when an employer bumps an old ad back to the top, evidence of the refresh. Datapika's Indeed and ZipRecruiter scan actor reads those pages and returns posted_at as a UTC timestamp such as 2026-07-07T06:55:00Z, plus a separate reposted_at whenever a listing was refreshed rather than newly created. It also pulls the employer's live ZipRecruiter posting total (593 for Paychex in the sample output) and a profile with industry, headcount, revenue band, headquarters, founding year, and a ZipRecruiter rating. Point it at a list of company names, optionally with websites for domain-level matching, and it reports has_active_postings and posted_last_48h_count per company; keyword mode takes up to 10 terms instead. Each company or job record costs $0.0005, and the hiring signals API use case page walks through the setup. Timestamp filtering replaces the board's own date facets, so a 48-hour window is true to the hour. - posted_at: exact UTC posting time read from the job page, not the search listing - reposted_at: present only when the employer bumped an existing ad, so fresh lists stay clean - Live ZipRecruiter total per employer (active_jobs_total_by_site), read from its employer page rather than a page-limited sample - Profile fields: zr_rating, zr_headquarters, zr_year_founded, zr_total_jobs alongside Indeed company_* keys - Company matching by website domain first, then normalized name with a 0.82 default threshold - recencyMode last48h or latest10, with hoursOld from 1 to 720 and up to 200 jobs per company per board ### How do you filter ZipRecruiter results by recency, remote, and job type? The input form is the same for every board, and all of it applies to ZipRecruiter. searchTerm takes one query, or searchTerms takes up to 5 that run in sequence and merge, with each row tagged by the term that surfaced it. location accepts a city, state, or country, and distance sets the search radius in miles with a default of 50. maxResults caps rows per board per term between 1 and 100, defaulting to 20, and offset skips ahead for pagination. isRemote restricts to remote roles, jobType picks one of fulltime, parttime, contract, internship, or temporary, hoursOld limits to postings inside a window, and easyApply keeps only listings that apply on the board itself. Because ZipRecruiter batches its intake, pair hoursOld of 48 with a scheduled daily run rather than chasing a 24-hour window. Advanced inputs cover descriptionFormat, enforceAnnualSalary, a custom userAgent, a caCert for enterprise proxies, and proxyConfiguration. If you want the same query answered by Indeed, LinkedIn, and Glassdoor in one deduplicated dataset, add them to sites; the full board list is at /scrape. - searchTerms: up to 5 queries per run, each row carries matched_search_term - location plus distance in miles (default 50) for metro-area targeting - maxResults 1 to 100 per board per term, default 20; offset for paging - isRemote, jobType, hoursOld, and easyApply all work on ZipRecruiter - enforceAnnualSalary normalizes hourly and monthly pay to yearly - userAgent, caCert, and proxyConfiguration for enterprise network setups ### How do you run the ZipRecruiter scraper from code or an AI agent? Three entry points share one actor and one dataset. In the Apify Console you fill the form, press Start, and download the Output tab as JSON, CSV, Excel, XML, or RSS. From code, POST the input to https://api.apify.com/v2/acts/openclawai~job-board-scraper/runs with your token, or call run-sync-get-dataset-items to get the rows back in one response; the Python client is pip install apify-client followed by client.actor("openclawai/job-board-scraper").call(run_input=...), and the Node client mirrors it. For agents, add https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper to Claude, ChatGPT, or Cursor and the assistant can ask for remote nursing jobs in Toronto posted this week and reason over typed rows instead of HTML. Zapier, Make, and n8n connect through the Apify app, and a Console schedule with a webhook turns one query into a daily feed. The actor holds a 5.0 rating from 3 reviews, and 381 distinct users ran it in the 30 days to August 29, 2026. - Console: form in, dataset out as JSON, CSV, Excel, XML, or RSS - REST: POST to /v2/acts/openclawai~job-board-scraper/runs, or run-sync-get-dataset-items for a single call - Python and Node clients: apify-client with actor("openclawai/job-board-scraper").call(...) - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper - Zapier, Make, n8n, schedules, and webhooks through the Apify platform - 2,471 total users, 381 active in the last 30 days, 27,458 runs as of August 29, 2026 ### Steps 1. Open https://apify.com/openclawai/job-board-scraper, enter a searchTerm such as registered nurse, set location to a US or Canadian city or state, and set sites to ["zip_recruiter"]. 2. Choose maxResults between 1 and 100, add hoursOld of 48 or more (ZipRecruiter ingests in daily batches), and toggle isRemote, jobType, or easyApply as needed; leave proxyConfiguration on residential. 3. Start the run. ZipRecruiter rows stream into the dataset after the browser warm-up, usually within 1 to 3 minutes; export JSON, CSV, or Excel from the Output tab or read the dataset through the API or MCP. 4. If you need exact posting times, repost flags, or a company's live ZipRecruiter job total, run https://apify.com/openclawai/indeed-ziprecruiter-scraper in companies mode with your target employer list. ### FAQ Q: Does ZipRecruiter have an official API? A: Not for searching jobs. ZipRecruiter's ZipSearch API, which let publishers pull listings, was deprecated on March 31, 2025, and as of July 2026 the publisher program has not reopened (Job Boardly). What remains is the Partner Platform Jobs API, which lets registered ATS partners create, update, retrieve, and close their own listings; access requires partner registration and an API key sent with every request (ZipRecruiter Partner Platform Documentation). Reading the public search pages is the only self-serve route to other employers' postings. Q: Why do ZipRecruiter rows appear later than Indeed rows in the same run? A: ZipRecruiter is fetched through a real browser session that has to pass an anti-bot warm-up before the first search, while Indeed and LinkedIn are read directly. In practice Indeed rows land 5 to 20 seconds after start and ZipRecruiter rows follow within 1 to 3 minutes. Results are pushed per board as each one finishes, so nothing waits on ZipRecruiter; you can read the early rows while its browser session is still working. Q: Can I scrape ZipRecruiter jobs outside the United States and Canada? A: No. ZipRecruiter's job search covers the US and Canada only, and the actor has no country parameter for it. For other markets, run the same query on Indeed or Glassdoor with countryIndeed set to uk, australia, germany, india, or another supported code, on LinkedIn which is global, on Naukri for India, or on Bayt for the Middle East. Cross-board duplicates collapse into one row tagged with each source, so mixing boards costs nothing extra in cleanup. Q: How fresh are the ZipRecruiter postings the scraper returns? A: ZipRecruiter ingests new listings in daily batches, so the newest job on the board is typically 12 to 48 hours old at scrape time. A 24-hour hoursOld window against ZipRecruiter alone will look sparse and is mostly driven by Indeed in a mixed run. Use 48 hours as the floor, schedule the actor daily, and if you need the true posting hour rather than the date, use the scan actor's posted_at field. Q: Does the scraper detect reposted ZipRecruiter jobs? A: The job board scraper returns date_posted as shown on the listing and deduplicates identical postings across boards, but it does not flag bumps. The Indeed and ZipRecruiter scan actor does: it keeps the original posting date and emits a separate reposted_at timestamp when an employer refreshed an existing ad, so a list filtered to the last 48 hours contains only genuinely new roles. That distinction matters most for hiring-velocity scoring and recruiter outreach. Q: Do I need my own proxies to scrape ZipRecruiter? A: No. The actor's proxyConfiguration defaults to Apify's residential proxy group, which is the right setting for ZipRecruiter because the board rate-limits datacenter IPs and the browser-based fetch has to clear an anti-bot warm-up first. Proxy traffic is billed as Apify platform usage rather than through the per-row price. Enterprise networks that route through their own gateway can supply a caCert and a custom userAgent instead. ## How to scrape Naukri job postings for the India market Canonical: https://datapika.com/scrape/naukri-jobs Updated: 2026-08-29 Naukri publishes no public developer API, so the practical route is the Datapika job board scraper on Apify. Set sites to naukri, enter a search term such as "data engineer", and each match lands as one row carrying four fields no other board supplies: skills, experience_range, vacancy_count, and work_from_home_type, plus company_rating and salary min, max, and currency. Naukri is fetched through a real browser, so a run typically takes 1 to 3 minutes. Pricing is $0.005 per job delivered, and the scraper shows 27,458 runs and a 5.0 rating from 3 reviews as of August 29, 2026. ### What fields does a Naukri job scraper return per posting? Every Naukri row uses the same flat schema as the other boards, then adds columns that only exist because Naukri exposes them. The standard part covers id, title, company, location, job_url, a job_url_direct apply link when the posting has one, date_posted, job_type, is_remote, and the full description in Markdown or HTML. Company context comes as company_rating, company_url, company_logo, and company_description where Naukri shows them. The Naukri-only part is the reason to scrape this board rather than rely on a global aggregator. skills holds the skills extracted from the posting, experience_range holds the years of experience the employer asks for, vacancy_count holds the number of open positions, and work_from_home_type holds Naukri's own hybrid, remote, or office label. Salary arrives as salary_min, salary_max, salary_currency, and salary_interval, with salary_source telling you whether the figures came from structured data or were parsed out of the description text. Each row also records matched_search_term and a scraped_at timestamp, so multi-term sweeps and repeat runs stay traceable. - skills: the skills extracted from the posting, populated on Naukri and no other board - experience_range: the years of experience Naukri lists as required, low end to high end - vacancy_count: open positions on the posting, a direct hiring-intensity signal - work_from_home_type: Naukri's WFH category, separate from the generic is_remote flag - company_rating: Naukri's employer score, carried on the same row as company - salary_min, salary_max, salary_currency, salary_interval, and salary_source when a range is posted ### What is different about scraping Naukri compared with Indeed or LinkedIn? Naukri is one of the boards the scraper reads through a real browser rather than a plain HTTP request, which is why the README puts it at 1 to 3 minutes per run while Indeed and LinkedIn land in 5 to 20 seconds. Rows stream in as each board finishes, so a mixed run shows faster boards first. Coverage is India only. The countryIndeed input does nothing for Naukri itself; it only matters if you add Indeed or Glassdoor to the same run, in which case set it to india so those boards return Indian inventory too. Naukri postings quote pay in rupees, so check salary_currency on each row before comparing against USD rows from other boards. enforceAnnualSalary converts monthly figures to yearly equivalents. maxResults caps at 100 per board per search term, searchTerms takes up to 5 queries, and offset paginates past the first page. Skills and experience are output columns, not input filters, so pull a wide set and filter after the run. - Browser-fetched: expect 1 to 3 minutes, versus 5 to 20 seconds for Indeed and LinkedIn - India coverage only; countryIndeed applies to Indeed and Glassdoor, not Naukri - salary_currency separates rupee ranges from USD rows in a multi-board dataset - enforceAnnualSalary normalizes monthly pay to a yearly figure - Up to 100 rows per search term per run, 5 terms per run, offset for pagination - Filter on skills and experience_range after the run; they are not search inputs ### How do recruiters and researchers use Naukri skills and experience data? The skills column turns a job search into a demand signal. Run the same 5 search terms every morning with hoursOld set to 24, count how often each skill appears across the fresh rows, and you have a daily view of which technologies Indian employers are asking for, broken down by city through the location field. experience_range adds a seniority axis without parsing description text: group rows by the minimum years requested and you can see whether a market is hiring juniors or leads. vacancy_count is the field sales and staffing teams watch. A posting with 15 open positions is a different lead from a single opening, and because company and company_rating sit on the same row, you can rank hiring companies by volume and by how their employees rate them. For salary research, salary_min and salary_max with salary_source set to direct_data give you employer-stated ranges you can trust more than figures parsed from prose. Export as CSV or Excel to pivot these columns, or push rows to Google Sheets or a CRM through Zapier, Make, or n8n. - Daily skill-demand tracking: hoursOld 24, count skills per run, split by location - Seniority mix from experience_range without reading descriptions - vacancy_count ranks companies by how many positions they are filling - company_rating next to company lets you qualify employers on the same row - salary_source direct_data marks employer-stated ranges for compensation studies - CSV, Excel, JSON, XML, or RSS export, plus Zapier, Make, and n8n connectors ### How do you call the Naukri scraper from code or an AI agent? From code, POST your input to https://api.apify.com/v2/acts/openclawai~job-board-scraper/runs with your Apify token and read the default dataset when the run finishes, or use the run-sync-get-dataset-items endpoint to get the rows back in one call. In Python, install apify-client and call client.actor("openclawai/job-board-scraper").call(run_input={"searchTerm": "DevOps engineer", "sites": ["naukri"], "maxResults": 30}), then list items from the dataset it returns. Node uses the same shape through apify-client on npm. For agents, add https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper to Claude, Cursor, or any MCP client, and the model can ask for Naukri postings and reason over the structured rows without touching HTML. Schedule the run in Apify Console, attach a webhook, and each finished run triggers your downstream processing. Naukri is one of 8 boards the same actor covers, and the setup guide for Claude Desktop and Cursor lives at datapika.com/actors/job-board-scraper. - REST: POST to acts/openclawai~job-board-scraper/runs, then read the default dataset - One-call option: run-sync-get-dataset-items returns the rows directly - Python and Node: apify-client with sites set to ["naukri"] - MCP: mcp.apify.com tool URL for Claude, ChatGPT, and Cursor - Schedules and webhooks in Apify Console for a daily India feed ### Steps 1. Open the job board scraper on Apify, enter a search term such as "backend developer" or up to 5 terms in searchTerms, and set sites to naukri (add indeed with countryIndeed india if you want a second Indian source). 2. Choose maxResults up to 100, optionally hoursOld for fresh postings only, and enforceAnnualSalary if you want yearly figures; then start the run and allow 1 to 3 minutes for the browser fetch. 3. Open the dataset, filter on skills, experience_range, vacancy_count, or salary_currency, and export as JSON, CSV, or Excel, or read the rows through the API or MCP. ### FAQ Q: Does Naukri have an official API? A: No. Naukri offers no developer portal, self-serve key, or public API documentation; the only integrations are commercial employer-side arrangements for posting jobs from an ATS and pulling applications back, negotiated with Naukri's team (JobsPipe, July 31, 2026). Datapika reads the public search pages instead, so there is no key or partner agreement, and you receive the same fields a visitor sees in the browser. Q: What does a Naukri scrape cost per job? A: The rate is $0.005 per job row delivered, so 1,000 Naukri postings cost $5 and a 30-row test run costs 15 cents. Charges apply only to rows that reach your dataset; a search that returns nothing costs nothing beyond Apify's small start fee. Apify's free plan includes $5 of monthly credit, which covers roughly 1,000 rows before you add a card. Q: Why does a Naukri run take longer than an Indeed run? A: Naukri, along with Glassdoor, ZipRecruiter, and Bayt, is fetched through a real browser with anti-bot warm-up, so the README lists 1 to 3 minutes for these boards versus 5 to 20 seconds for Indeed and LinkedIn. Since each board's rows are pushed the moment that board finishes, a mixed run shows Indeed rows first and Naukri rows when its browser session completes. Q: Can I filter Naukri results by skills or years of experience? A: Not as inputs. The scraper accepts searchTerm, location, jobType, isRemote, hoursOld, and easyApply as run-level filters, and returns skills and experience_range as output columns on each Naukri row. The reliable pattern is to set maxResults to 100, pull the widest set for your term, and filter the dataset afterwards in a spreadsheet, SQL, or your own code. Q: Are Naukri salaries returned in rupees? A: Each row carries salary_min, salary_max, salary_currency, and salary_interval. Naukri postings quote pay in rupees, and salary_currency records the currency the board reports, so check it before mixing rows with USD boards. salary_source tells you whether the numbers were board-provided (direct_data) or parsed from the description. Turn on enforceAnnualSalary to convert monthly figures to yearly equivalents before comparing across boards. Q: How many Naukri jobs can one run return? A: maxResults caps at 100 per board per search term and searchTerms accepts 5 queries, so a Naukri-only run tops out at 500 rows. Use offset to skip rows you already have on a follow-up run, or split the sweep by city and job type. Deduplication removes repeats across boards, so a posting that also appears on Indeed India counts once. ## How to scrape Bayt.com job postings across the Gulf and MENA Canonical: https://datapika.com/scrape/bayt-jobs Updated: 2026-08-29 Bayt.com has no public jobs API, so the practical route is Datapika's job board scraper on Apify with bayt in the sites list. Bayt is a keyword-only board: the scraper honors searchTerm and up to 5 searchTerms, and ignores location, hoursOld, jobType and remote filters, so the city or country belongs inside the query itself. Bayt pages are rendered in a real browser, which puts a typical run at 1 to 3 minutes, and each posting lands as one row with title, company, location, job URL and the full description. Pricing is $0.005 per job delivered, about $5 per 1,000. ### Which countries and roles does a Bayt scraper cover? Bayt.com is the regional board the scraper uses for the Middle East, and its own app listing describes 58 million job seekers and 60,000+ employers across the UAE, Saudi Arabia, Qatar, Kuwait, Egypt, Jordan, Lebanon and the wider MENA region since 2000. That makes it the natural source for Gulf roles that never reach LinkedIn or Indeed: Saudi construction supervisors, Dubai retail managers, Qatar hospitality staff, Egyptian accountants posted by a Riyadh employer. Bayt operates in English and Arabic, and the scraper stores the description exactly as the employer wrote it, so a single dataset can mix Latin and Arabic script and your downstream filters should tolerate both. Because Bayt has no location input on the scraper side, coverage is steered purely by the words you search. A query such as "HSE officer Saudi Arabia" narrows to one country, while "HSE officer" alone returns the board's regional ranking. Visa sponsorship and nationality preferences are not separate fields; when an employer states them, they appear in the description text and can be extracted with a simple keyword pass. - Regional scope per Bayt's own listing: UAE, Saudi Arabia, Qatar, Kuwait, Egypt, Jordan, Lebanon and the wider MENA region - 58 million job seekers and 60,000+ employers on the platform, operating since 2000 - Postings can be in English, Arabic, or a mix; descriptions arrive unchanged in Markdown or HTML - No location input applies to Bayt, so put the country or city inside the search term - Visa, sponsorship and nationality requirements live in the description text, not in dedicated columns - Pair Bayt with Naukri in one run to cover the Gulf and India expat corridor together ### What fields does each Bayt job row contain? Every Bayt posting becomes one dataset item that shares the same column layout as the other boards, with site set to bayt so rows can be filtered after a multi-board sweep. The reliably populated core is id, title, company, location, job_url, date_posted, description, search_term, matched_search_term and scraped_at. The description is the full posting body in Markdown by default or HTML if you set descriptionFormat, and emails found inside it are copied into the emails array, which matters whenever a Bayt employer asks applicants to write in directly. Salary columns are filled when the posting states a figure: salary_source reports whether the number came from structured board data or was parsed out of the description, alongside salary_min, salary_max, salary_currency and salary_interval, and enforceAnnualSalary converts monthly AED or SAR figures to yearly equivalents. Columns that the README ties to other boards stay null on Bayt rows: job_level is LinkedIn only, company_country is Indeed only, and skills, experience_range, vacancy_count and work_from_home_type are Naukri only. The table below marks what to expect column by column. - Core on every row: id, title, company, location, job_url, site, date_posted, description, scraped_at - matched_search_term shows which of your up-to-5 queries surfaced the posting - emails array captures contact addresses embedded in the posting text - Salary quartet (min, max, currency, interval) plus salary_source when the employer states pay - enforceAnnualSalary turns monthly Gulf salaries into yearly figures for comparison - job_level, company_country and the four Naukri columns are null on Bayt rows ### What are Bayt's quirks compared with the other boards? Bayt behaves differently from the URL-parameter boards, and knowing the differences avoids empty or overpriced runs. First, it is the board where searchTerm is the sole filter: location, distance, isRemote, jobType, hoursOld, easyApply and offset are all ignored for Bayt, so a run with hoursOld: 24 still returns Bayt's default ordering rather than fresh postings. Second, Bayt is fetched through a stealth browser rather than a plain HTTP request, which is why its rows land after 1 to 3 minutes while Indeed and LinkedIn stream in within 5 to 20 seconds; runs default to 4 GB of memory to give the browser headroom. Third, freshness has to be handled on your side: keep the date_posted column, store the id of every row, and diff successive runs to isolate new postings. Fourth, the language mix means a search term in English may not surface a posting written only in Arabic, so bilingual sweeps need both spellings as separate entries in searchTerms. - searchTerm is the only filter Bayt honors; every other input is ignored for this board - Browser-rendered board: expect 1 to 3 minutes, versus 5 to 20 seconds for Indeed and LinkedIn - Runs default to 4 GB memory so the browser-based boards have headroom - No hoursOld effect on Bayt, so diff id values between scheduled runs to find new postings - Add an Arabic spelling of the role as a second search term to catch Arabic-only postings - Cross-board deduplication merges a Bayt posting that also appears on LinkedIn into one row ### How many Bayt jobs can one run return, and how fast? The ceiling is set by two inputs. maxResults caps each board at 100 rows per search term, and searchTerms accepts up to 5 queries, so a Bayt-only run tops out at 500 rows. Because Bayt ignores offset, paging beyond that means changing the query: split a broad title into per-country variants such as "procurement manager UAE", "procurement manager Saudi Arabia" and "procurement manager Qatar", or by seniority. Results stream to the dataset as each board finishes, so in a mixed run you can already read Indeed rows while the Bayt browser session is still loading; a verified default-style run produced its first dataset rows 3 seconds after the container started. For a Bayt-only run the wait is the browser warm-up plus page loads, typically 1 to 3 minutes end to end. Keep maxResults at 20, the default, until the query returns the roles you want. The actor has 27,458 runs and 2,471 users on Apify as of August 29, 2026, with 381 active in the last 30 days. - maxResults: 1 to 100 per board per search term; searchTerms: up to 5 queries - Bayt-only maximum per run: 500 rows; widen coverage by varying the country in the term - offset is ignored on Bayt, so pagination happens through query variants - First rows in a mixed run appear about 3 seconds after container start; Bayt rows follow in 1 to 3 minutes - 27,458 runs, 2,471 users, 381 active in the last 30 days, 5.0 rating from 3 reviews (Apify, August 29, 2026) - Default maxResults of 20 keeps test runs small while you tune the search term ### How do you call the Bayt scraper from code or an AI agent? The input is a single JSON object, and the Bayt-only version is short: searchTerm or searchTerms, sites set to ["bayt"], and maxResults. From the Apify Console you paste that into the form, run, and download the dataset as JSON, CSV, Excel, XML, HTML or RSS. From code, POST the same object to the run-sync-get-dataset-items endpoint for openclawai~job-board-scraper with your Apify token and the response body is the array of Bayt rows; the Python client wraps this as client.actor("openclawai/job-board-scraper").call(run_input=...), and the Node client mirrors it with apify-client. For agents, add https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper to Claude Desktop, Cursor or any MCP client and the assistant can answer prompts like "which Dubai employers are hiring civil engineers on Bayt this week" by calling the tool and reasoning over the rows. Zapier, Make and n8n connect through the Apify app, and a daily schedule with a webhook turns the run into a standing Gulf job feed. The same run can sweep LinkedIn, Indeed, Glassdoor, ZipRecruiter and Naukri alongside Bayt; see the full board list at /scrape. - Minimal input: {"searchTerms": ["nurse Riyadh", "nurse Dubai"], "sites": ["bayt"], "maxResults": 50} - REST: POST to run-sync-get-dataset-items for openclawai~job-board-scraper and parse the returned array - Python: pip install apify-client, then client.actor("openclawai/job-board-scraper").call(run_input=...) - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper for Claude, ChatGPT or Cursor - Zapier, Make and n8n push new rows to Sheets, Airtable, Slack or a CRM on a schedule ### Steps 1. Open https://apify.com/openclawai/job-board-scraper, set sites to ["bayt"], and enter a searchTerm that includes the country or city, for example "finance manager Dubai". 2. Add up to 4 more variants in searchTerms (other Gulf countries, an Arabic spelling of the title) and set maxResults between 20 and 100 per term. 3. Start the run and wait 1 to 3 minutes for the browser session to finish; rows appear in the dataset with site = bayt and a matched_search_term per row. 4. Export as JSON or CSV, or fetch the rows through the Apify API, then schedule the run daily and diff id values to isolate new postings. ### FAQ Q: Does Bayt have an official API? A: No public one. Bayt.com offers no developer program for reading job listings, and third-party scraper listings on Apify state plainly that no official public API exists. Employers post through Bayt's own dashboard, so reading listings programmatically means scraping the public search pages, which is what this actor does without a login, an API key or a partner agreement. The rows you get carry the same title, company, location and description you would see in a browser. Q: Can I filter Bayt jobs by location, date, or remote status? A: Not through the scraper inputs. Bayt honors only searchTerm and searchTerms; location, distance, hoursOld, jobType, isRemote, easyApply and offset are ignored for this board. Put the country or city in the query, for example "electrical engineer Kuwait", and handle freshness after the fact by comparing date_posted and id across scheduled runs. The same inputs still work on LinkedIn, Indeed and Glassdoor within a mixed run. Q: Does the scraper return Arabic postings from Bayt? A: It returns whatever the posting contains. Bayt operates in English and Arabic, and the description column stores the employer's text as written, so rows can be in English, Arabic or both. Matching is keyword based, so an English search term may not surface a listing written only in Arabic; add the Arabic job title as a separate entry in searchTerms and both sets of rows arrive in the same dataset with matched_search_term telling them apart. Q: How much does it cost to scrape 1,000 Bayt jobs? A: $5, at $0.005 per job row delivered to your dataset, with no monthly fee. A Bayt-only run is capped at 500 rows (100 per term times 5 terms), so the largest single run costs $2.50, and a run that returns nothing costs only Apify's small start fee. Apify's free-plan monthly credit covers roughly 1,000 Bayt postings before you pay anything. Q: Are visa, sponsorship, or nationality requirements available as fields? A: Not as dedicated columns. The dataset schema has no visa or nationality field for any board, so when a Gulf employer specifies sponsorship, a preferred nationality, or a required residency, the statement sits inside the description text. Because the description arrives complete in Markdown, a keyword match for terms like "visa", "sponsorship" or a nationality name, or a short LLM pass over the text, extracts it reliably after the run. Q: Why do Bayt rows take longer to appear than Indeed rows? A: Bayt is one of the boards fetched through a real browser with anti-bot warm-up, along with Glassdoor, ZipRecruiter and Naukri, so its pages typically take 1 to 3 minutes to load and parse. Indeed and LinkedIn are read over plain requests and land within 5 to 20 seconds. Since v1.0.49 each board pushes its rows the moment it finishes, so a mixed run lets you read fast boards while Bayt is still rendering. ## How to track competitor hiring with a scheduled weekly scan Canonical: https://datapika.com/use-cases/track-competitor-hiring Updated: 2026-08-29 To track competitor hiring, put your competitors' names into companies mode on Datapika's Indeed and ZipRecruiter scraper and schedule the run once a week on Apify. Each company comes back with a yes or no has_active_postings signal, verified posting counts split by board, the live ZipRecruiter job total (593 for the sample company in the actor README), a firmographic profile, and every job posted inside a window you can stretch to 720 hours. Records cost $0.0005 each, company and job rows alike, so the scan is cheap and the week-over-week diff of totals and titles is the intelligence. ### What does a weekly scan of a named competitor return? One company record per name, plus one flat row per job when emitJobRecords is on. The record opens with company_query and matched_company_name, so you can see which employer the boards resolved your input to. has_active_postings answers the yes or no question, active_jobs_by_site splits the employer-verified postings between Indeed and ZipRecruiter, and active_jobs_total_by_site.zip_recruiter reads the company's full open count from its ZipRecruiter page rather than the sample this run pulled. Indeed exposes no such total, so that slot is null. posted_last_48h_count tells you how many postings landed inside your recency window, which you set with hoursOld. Under profile sit industry, employee count, revenue band, headquarters, founding year and employee rating, with Indeed keys and ZipRecruiter keys side by side. latest_jobs nests the postings with title, location, salary fields, job_type, is_remote, job_level, posted_at and reposted_at. sites_with_errors names any board that failed for that company, and scraped_at stamps the run in UTC so each week's dataset sorts cleanly against the last. - Sample record in the README: 17 Indeed and 21 ZipRecruiter postings found, live ZipRecruiter total 593, 17 posted inside the window - Profile keys include company_industry, company_num_employees, company_revenue, zr_rating, zr_headquarters and zr_year_founded - Each nested job also lands as a flat record_type: job row for spreadsheet-friendly exports - posted_at is hour-precise on ZipRecruiter and day-precise on Indeed - sites_with_errors lets you rerun only the companies where one board failed ### How do you schedule a weekly competitor scan on Apify? Start with the companies input. Write each competitor as name and website, for example Kelly Services | kellyservices.com, because domain matching stops a generic name from pulling in unrelated employers; the normalized-name fallback is controlled by matchThreshold, default 0.82. Leave scanMode on companies and recencyMode on last48h, then set hoursOld to 168 so a run captures everything posted since the previous one; the field accepts 1 to 720 hours. maxJobsPerCompany caps the rows per company per board at 50 by default and 200 at most, which is where you control dataset size. Keep both sites enabled unless one board is irrelevant to your market, set countryIndeed if your competitors hire outside the US, and keep the residential proxy, since ZipRecruiter rate-limits datacenter IPs. Save the input, then create an Apify schedule with a weekly cron and attach a webhook that fires on run completion. A ten-company run with the default window finishes in a few minutes; split lists of several hundred names across schedules so each run stays inside the 60-minute default timeout. - companies: Name | website per line, or {name, website} objects through the API - hoursOld 168 with recencyMode last48h covers a full week between runs; latest10 returns the 10 newest postings regardless of age - maxJobsPerCompany: 50 default, 200 maximum, per company per board - countryIndeed accepts usa, uk, canada and others; ZipRecruiter covers the US and Canada - Turn includeDescription on only if you need full job text, since it adds latency on Indeed - Weekly schedule plus completion webhook makes the run hands-off after the first setup ### How do you diff two weekly runs to spot a hiring shift? Every run writes its own dataset, so the comparison is a join on company_query between this week's export and last week's. Three deltas carry most of the meaning. The change in active_jobs_total_by_site.zip_recruiter is the volume signal: a jump means new requisitions were approved, a drop means roles closed or a freeze started. posted_last_48h_count, with the window set to 168 hours, counts only what appeared since your last run, so it doubles as a velocity number without any date math on your side. The third delta is in the job rows: new title strings that did not exist last week, new locations, and new values in job_level. Filter out rows that carry reposted_at before counting, because a bumped ad is not a new opening and would inflate the week's total. Keep scraped_at on every row and append each run to one table; after a quarter you have a dated series per competitor that a pivot or a notebook can chart directly. - Join key: company_query, which echoes the exact name you submitted - Volume delta: active_jobs_total_by_site.zip_recruiter this week minus last week - Velocity: posted_last_48h_count with hoursOld at 168, no date parsing needed - Novelty: titles, locations and job_level values absent from the previous run - Exclude rows with reposted_at from new-opening counts; track them separately as hard-to-fill roles - Use the companies and jobs dataset views for pre-flattened tables before joining ### Which week-over-week changes are worth acting on? Counts on their own tell you that something moved; the job rows tell you what. A competitor whose first posting in a new country or metro shows up in location is opening a market, and it usually appears there before any announcement. A cluster of engineering titles that share a product name points to a build in progress, while a batch of account executive or customer success roles points to a go-to-market push. job_level shifting upward, with director and VP titles appearing where individual contributor roles used to be, suggests a new function being stood up. salary_min and salary_max on each row put a number on how aggressively they are bidding for talent against your own bands. On the negative side, a falling live total with no fresh postings for two consecutive runs reads as a freeze, and reposted_at stacking up on the same title reads as a role they cannot fill. The profile block adds the employee rating, which moves slowly but frames retention. - New location value for a competitor: market entry, so check which function leads the titles - Cluster of titles sharing a product or platform name: a build underway - Sales and customer success roles appearing together: a go-to-market push - job_level moving to director and above: a new function or leadership layer - Live total falling for two runs with zero fresh postings: freeze or restructuring - reposted_at repeating on one title: a role they are struggling to fill ### How do you get the weekly diff into a dashboard, CRM or AI agent? The run finishes, the webhook fires, and the dataset is ready. The REST endpoint for dataset items returns JSON, CSV or XLSX, and the jobs and companies views give you two pre-flattened tables ready for a warehouse or a Google Sheet. The Python and JavaScript clients wrap the same call for teams that want to run the diff in a notebook or a cron job. On the no-code side, the actor's Integrations tab on Apify lists Zapier, Make, Google Sheets, Airtable and Slack, so a Slack post per competitor whose live total jumped is a few clicks. For agents, the scraper is exposed as an MCP tool at mcp.apify.com, which lets a Claude or ChatGPT based assistant answer a question like which of these competitors opened new roles this week by running the scan and reading the rows. Pay-per-call flows over X402 and MPP are supported for agents without an Apify subscription. Combine it with the other job boards in the Datapika catalog at /scrape when a competitor hires mostly through LinkedIn or Glassdoor. - REST: dataset items endpoint with format=json, csv or xlsx, plus the jobs and companies views - Python: pip install apify-client, then actor('openclawai/indeed-ziprecruiter-scraper').call(run_input={...}) - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/indeed-ziprecruiter-scraper - No-code: Zapier, Make, Google Sheets, Airtable and Slack from the Integrations tab - Webhook on run completion triggers the diff step or the Slack alert ### Steps 1. Open https://apify.com/openclawai/indeed-ziprecruiter-scraper, keep scanMode on companies, and paste your competitor list one per line as Name | website so employer matching resolves to the right company. 2. Set hoursOld to 168 with recencyMode last48h so each run covers the week since the last one, leave maxJobsPerCompany at 50 or raise it toward 200 for heavy hirers, set countryIndeed for non-US competitors, and keep the residential proxy. 3. Run once to confirm matched_company_name is correct for every competitor, then create an Apify schedule with a weekly cron and a webhook on completion. 4. After each run, pull the companies and jobs dataset views, join on company_query to the previous week, and act on rises in the live ZipRecruiter total and on new titles and locations, excluding rows with reposted_at. ### FAQ Q: How often should I scan competitor job postings? A: Weekly is the practical cadence for competitor tracking. Indeed timestamps postings the same day, but ZipRecruiter ingests in daily batches, so its newest postings are typically 12 to 48 hours old; a 168-hour window on a weekly schedule catches everything from both boards without gaps. Daily runs make sense for a shortlist of two or three direct rivals where the extra rows are cheap. Whatever the cadence, keep hoursOld equal to the interval between runs. Q: Does Indeed have an official API for tracking competitor job postings? A: No. Indeed retired its Publisher API, the read-side feed developers used to pull listings, in 2023, and the Publisher Program itself closed to new sign-ups in October 2022 with no reopening announced (Job Boardly, 2026; JobsPipe, 2026). What remains, Indeed Apply, Sponsored Jobs and the Job Sync API, is employer-side and gated behind partner approval at partners.indeed.com. There is no self-serve endpoint that returns Indeed search results or company pages, which is the gap this scan fills. Q: Does ZipRecruiter have an official API I could use instead? A: Not for reading postings. ZipRecruiter shut down its ZipSearch publisher API on March 31, 2025, ending the program that let sites query its listings (Job Board Secrets, 2025; Job Boardly, 2026). The partner API it still runs is for job boards and publishers that syndicate listings to ZipRecruiter, not for querying jobs or company pages on demand. Neither path exposes an employer's live job total or hour-precise posting timestamps, which are the two fields the weekly diff depends on. Q: How much does a weekly competitor scan cost? A: Every company record and every job row is one result event at $0.0005, so a scan returning 50 company records and 950 job rows costs $0.50 in results, plus a fraction of a cent for the run start. Apify platform usage (compute and proxy traffic) is billed separately and is typically a few cents for a run that size; every account includes free monthly platform credit. To keep a large list cheap, lower maxJobsPerCompany, turn emitJobRecords off, or use profiles mode when you only need the live total. Q: How do I stop a competitor with a common name from matching the wrong employer? A: Put the website after the name with a pipe between them, such as Kelly Services | kellyservices.com, or pass {name, website} objects through the API. Jobs are matched by the employer's domain first and only fall back to a normalized-name comparison governed by matchThreshold, which defaults to 0.82 on a 0.5 to 1 scale. Check matched_company_name in the first run for every competitor; if a name resolved wrongly, raise the threshold or add the domain and rerun that company alone. Q: Can I track competitors that hire outside the United States? A: Partly. Indeed supports country selection through countryIndeed, with values such as usa, uk and canada, so a UK or Canadian competitor list works by switching that field and setting location to match. ZipRecruiter covers the US and Canada only, so for other markets the ZipRecruiter live total will be empty and the signal rests on Indeed counts and job rows. For competitors hiring mainly in India or the Gulf, the Naukri and Bayt scrapers in the catalog cover those boards. ## Search remote jobs on LinkedIn, Indeed, Glassdoor, ZipRecruiter, Naukri and Bayt in one sweep Canonical: https://datapika.com/use-cases/remote-job-search Updated: 2026-08-29 A remote job search that covers every major board takes one run of the Datapika job board scraper: set isRemote to true, enter up to 5 search terms, and it sweeps LinkedIn, Indeed, Glassdoor, ZipRecruiter, Naukri and Bayt, then removes cross-posted duplicates so each opening appears once. Every row carries title, company, location, salary range, is_remote flag and a direct apply URL. Indeed and LinkedIn rows land within 5 to 20 seconds. It costs $0.005 per job delivered, and 2,471 people have used it across 27,458 runs with a 5.0 rating. ### How do I search remote jobs on every job board at once? Start with the isRemote switch, which asks each board for remote listings only, and put your role names in searchTerms, up to 5 per run, so titles like backend engineer, platform engineer and site reliability engineer merge into one dataset with a matched_search_term column showing which query found each row. Leave location empty or set it to Remote; giving a city instead, with the default 50 mile distance, returns remote listings the board associates with that area, which helps when an employer hires remotely but only inside one country. Pick boards in sites: linkedin, indeed, glassdoor, zip_recruiter, bayt and naukri run by default. Set maxResults between 1 and 100 per board per term, and add jobType if you only want fulltime or contract work. Because the same posting often appears on LinkedIn, Indeed and Glassdoor together, the run removes duplicates before rows reach the dataset, so a six-board sweep reads like one board with far more listings. - isRemote: true applies the remote filter on every board that supports it in a single run - searchTerms accepts up to 5 titles; each row records matched_search_term - location can be blank, Remote, or a city with the 50 mile default distance for region-bound remote roles - maxResults ranges from 1 to 100 per board per search term - jobType narrows to fulltime, parttime, contract, internship or temporary - Default sites are linkedin, indeed, glassdoor, zip_recruiter, bayt and naukri ### Which job boards cover remote roles, and where does each one fall short? Each board behaves differently for remote searches. LinkedIn is global and accepts isRemote, but it rate-limits at roughly 100 results per IP, so the residential proxy default matters on large runs; enable linkedinFetchDescription when you need full descriptions and direct apply links from it. Indeed is the most dependable source and takes a countryIndeed code (usa by default, with uk, canada, australia, germany, france, india and others), so run it once per country to map where a remote role is actually open. Glassdoor needs the same countryIndeed setting for country targeting. ZipRecruiter only covers the US and Canada. Naukri serves India and returns a work_from_home_type field describing the work-from-home arrangement, plus skills and experience_range. Bayt covers the Middle East but accepts only searchTerm, so isRemote does not apply there and you filter its rows by description afterwards. Google Jobs and BDJobs stay selectable but currently return no results. The full list of boards, including the ones outside this search, is in the catalog at /scrape. - LinkedIn: global coverage, about 100 results per IP, residential proxy on by default - Indeed and Glassdoor: set countryIndeed (usa, uk, canada, australia, germany, france, india and more) - ZipRecruiter: US and Canada listings only - Naukri: India, with work_from_home_type, skills and experience_range fields - Bayt: Middle East, searchTerm only, so apply the remote filter on your side - Google Jobs and BDJobs: selectable but returning no results at the moment ### How does deduplication keep one row per remote job? Remote roles are heavily cross-posted: an employer that hires anywhere tends to publish on LinkedIn, Indeed and Glassdoor at once, and a naive multi-board pull repeats every row. The scraper collapses those duplicates before pushing to the dataset, so you review each opening once, and the site column tells you which board supplied the surviving copy. That is also why a run usually returns fewer rows than maxResults multiplied by boards and terms, and why the theoretical ceiling of 100 results x 6 default boards x 5 terms, 3,000 rows, is rarely reached. Billing follows the deduplicated count, not the ceiling. For your own tracker, key on the id field, which carries the board's job identifier, and drop any id you have already seen; a weekly run then yields only openings that are new to you. Use offset to page past the first results when a single title on one board has more than 100 remote matches. - Cross-board duplicates are removed automatically; each opening is one row - site names the board the row came from; id is that board's job identifier - Ceiling per run: 100 results x 6 default boards x 5 terms = 3,000 rows, with dedup usually returning fewer - You are billed for delivered rows only, never the theoretical maximum - Keep a seen-id list in your tracker to dedupe across repeat runs - offset skips the first N results for pagination ### What fields help a job seeker decide whether to apply? Every row is built to answer the questions you ask before spending an hour on an application. salary_min, salary_max, salary_currency and salary_interval come straight from the board when available, and salary_source tells you whether the figure was board-provided (direct_data) or parsed from the description text. Set enforceAnnualSalary to true and hourly and monthly rates convert to yearly equivalents, so a mixed remote list sorts on one column. company_rating and company_reviews_count let you screen employers, company_num_employees and company_industry show whether you are looking at a 50-person startup or a bank, and job_level on LinkedIn rows flags seniority. job_url_direct is the employer's own application page where the board exposes it, which avoids creating yet another job-board account. The description arrives as Markdown by default, or HTML if you prefer, so keyword checks for US only or EU hours are a regex away, and any emails found in the description are pulled into their own field. - salary_min, salary_max, salary_currency, salary_interval, plus salary_source for provenance - enforceAnnualSalary converts hourly and monthly pay to yearly figures - company_rating, company_reviews_count, company_num_employees, company_industry for employer screening - job_url_direct points at the employer's apply page when the board exposes it - description in Markdown or HTML; emails extracted into their own field - job_level (LinkedIn) and job_function for seniority and role category ### How do I turn this into a repeatable job-seeker workflow? The job-seeker loop is: run, dedupe against what you have seen, review, apply, repeat. On Apify, save the input once and attach a schedule so the sweep runs on the mornings you choose; each run produces a dataset you can download as JSON, CSV or Excel, or read through the API. For a no-code pipeline, the Apify app for Zapier, Make and n8n can append new rows to Google Sheets or Airtable and post a summary to Slack when a run finishes, and a webhook on run completion can trigger any script you already have. Keep maxResults low while you tune search terms, then raise it, and cap the run's total charge in Console so an over-broad query cannot surprise you. If you work with an AI assistant, add the MCP server URL for the actor and ask it in plain language for remote roles matching your profile; it receives the same structured rows. For a strictly recent feed, pair this setup with the hoursOld filter described on the fresh-jobs page. - Schedule the saved input in Apify Console to run on your chosen mornings - Export JSON, CSV or Excel, or pull rows through the API - Zapier, Make and n8n connectors push new rows to Google Sheets, Airtable or Slack - Webhook on run completion to trigger your own scoring or tracker script - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper - Cap the run's total charge in Console and start with a small maxResults while tuning ### Steps 1. Open https://apify.com/openclawai/job-board-scraper, enter up to 5 searchTerms for the roles you want, set isRemote to true, and leave location blank or set it to Remote (use a city plus countryIndeed if the role must sit within one country). 2. Keep the default boards (linkedin, indeed, glassdoor, zip_recruiter, bayt, naukri), pick a jobType if you need one, turn on enforceAnnualSalary, and start with maxResults at 20 to check the output shape. 3. Start the run. Indeed and LinkedIn rows land within 5 to 20 seconds and the browser-fetched boards follow within 1 to 3 minutes; sort the dataset by salary and company_rating to build your shortlist. 4. Export as CSV or JSON, or connect Zapier, Make or n8n to append new ids to your tracker, then schedule the run so future sweeps only surface openings you have not seen. ### FAQ Q: Does LinkedIn or Indeed have an official API for remote job search? A: No self-serve option exists on either. LinkedIn's only jobs API, the Job Posting API, pushes listings onto LinkedIn rather than reading them, is restricted to approved Talent Solutions partners, and its documentation states LinkedIn is currently not accepting new partnerships (Microsoft Learn, updated June 2026). Indeed closed its Publisher Program to new publishers in October 2022 (Job Boardly) and shut the Publisher API in 2023, with remaining partner access granted through sales-led approval (JobsPipe, July 2026). The scraper reads public search pages instead. Q: Why does Indeed return fewer remote jobs when I add hoursOld? A: Indeed's search cannot combine hoursOld with isRemote, jobType or easyApply in one query, so a run that sets both returns fewer Indeed rows than you expect. For a remote sweep that includes Indeed, leave hoursOld empty and filter on date_posted in your own tracker instead. LinkedIn has a similar restriction between hoursOld and easyApply. No such restriction is documented for the other boards. Q: How many remote jobs can one run return? A: maxResults caps at 100 per board per search term and searchTerms at 5, so the six default boards give a ceiling of 3,000 rows, or 4,000 if you also select Google Jobs and BDJobs, which currently return nothing. Deduplication across boards usually brings the actual count well under the ceiling. To go further, paginate with offset or split the run by country, job type or search term. Q: How much does a remote job search across all boards cost? A: Pricing is $0.005 per job delivered, which works out to $5 for 1,000 rows, and you pay only for rows that land in the dataset after deduplication. A run that finds nothing costs only the small Apify start fee. Apify's free plan includes $5 of monthly credit, roughly 1,000 jobs, so a weekly single-term sweep at 20 results across the six default boards fits inside it. Q: Does the remote filter work on every board? A: The scraper passes isRemote to every board except Bayt, which accepts only searchTerm, so Bayt rows arrive unfiltered; check is_remote and the description before applying. Naukri additionally returns work_from_home_type, which describes the work-from-home arrangement on each listing. On any board, some employers tag a role remote but restrict it to one country, so read the location field and scan the description for phrases like US only or EU hours. Q: Can I get new remote jobs sent to Google Sheets or Slack automatically? A: Yes. Schedule the run in Apify Console, then use the Apify app in Zapier, Make or n8n to append each new dataset to Google Sheets or Airtable and post a Slack message when the run finishes. Filter on the id field so the sheet only receives openings you have not seen before. A webhook on run completion works the same way for custom scripts or a personal tracker. ## Build a live salary dataset from job postings across 6 boards Canonical: https://datapika.com/use-cases/salary-market-research Updated: 2026-08-29 Salary market research from job postings starts with structured pay fields, and the Datapika job board scraper returns four on every row that states pay: salary_min, salary_max, salary_currency and salary_interval, plus a salary_source flag that says whether the board supplied the figure or it was parsed from the description. Switch on annual normalization and hourly rates are multiplied by 2,080 and monthly pay by 12, so contract and salaried roles share one column. One run sweeps LinkedIn, Indeed, Glassdoor, ZipRecruiter, Naukri and Bayt, deduplicates across boards, and costs $0.005 per job, so a 1,000-row study is about $5. ### Which salary fields does a scraped job posting include? Five pay fields sit on each row under the same names, whichever board produced it. salary_min and salary_max hold the numeric range as the posting states it, salary_currency holds the currency code such as USD or GBP, and salary_interval records the original pay period as yearly, monthly or hourly. The fifth field, salary_source, is the one to read first: direct_data means the board exposed the figure as structured data, while description means the number was extracted from the posting text. Structured values are the safer input for a median; parsed values catch postings that only mention pay in a paragraph, but they deserve a spot check before they enter a benchmark. Around the pay fields sit the segmentation columns that explain variance: job_type, is_remote, location, job_level on LinkedIn, and company_num_employees, company_revenue, company_industry and company_rating where the board publishes an employer profile. The sample Indeed row in the listing shows the shape: 120,000 to 185,000 USD yearly from direct_data, with a 3.9 employer rating and 21,432 reviews on the same row. - salary_min and salary_max: numeric bounds exactly as the posting states them, null when no pay is published - salary_currency: the currency code, so UK, Indian and Gulf rows never blend into USD averages by accident - salary_interval: yearly, monthly or hourly as published, which tells you which rows need converting - salary_source: direct_data from the board or description parsed from text, your first quality filter - Segmentation on the same row: job_type, is_remote, job_level, employer size, revenue, industry and rating - Descriptions ship as Markdown or HTML, so you can re-parse pay language yourself when a posting is ambiguous ### How do you normalize hourly, monthly and yearly salaries into one column? Postings quote pay in whatever unit the employer typed, and an hourly contract sitting next to an annual salaried range breaks any median you compute. The scraper has one input for this, enforceAnnualSalary. When it is on, hourly figures are multiplied by 2,080 (40 hours across 52 weeks) and monthly figures by 12, so a 45 per hour contract lands at 93,600 and a 7,500 per month role at 90,000, in the same column as a posted 95,000 salary. Leave it off when hourly and salaried markets should be studied separately, for example nursing or trades roles where hourly is the native unit; salary_interval then tells your own code which rows to convert. Normalization never converts currency. A run with countryIndeed set to uk returns GBP rows and a Naukri sweep returns Indian postings, so group by salary_currency before ranking anything, or apply your own exchange rates at analysis time. Keep salary_source in view as well: a description-parsed hourly number multiplied by 2,080 amplifies any parsing error by the same factor. - enforceAnnualSalary: hourly x 2,080, monthly x 12, yearly left unchanged - 45 per hour becomes 93,600 a year; 7,500 per month becomes 90,000 - Off by default, so hourly-native markets keep their original unit unless you ask - Currency is never converted: group by salary_currency or apply your own rates - Parsed hourly values are the rows to spot check, because the 2,080 multiplier scales any error ### Which job boards publish salary data, and what does each one add? Salary coverage depends on the board and on the employer, so the practical rule is to sweep every board that serves your market and let deduplication merge the overlaps; per-board detail pages are listed at /scrape. Indeed is the most reliable source and carries the richest employer profile, including employee count, revenue band and company country, with countryIndeed choosing the national site. Glassdoor uses the same countryIndeed setting for country targeting and is fetched through a real browser, so its rows arrive after Indeed's. LinkedIn adds job_level for seniority cuts but rate-limits at roughly 100 results per IP, so keep residential proxies on and enable linkedinFetchDescription when you need description-parsed pay. ZipRecruiter covers the US and Canada only. Naukri returns experience_range, skills and company_rating, which makes it the board to use for pay-by-experience curves in India, while Bayt covers the Middle East and accepts only the search term, with no hours or job type filters. Google Jobs and BDJobs stay selectable but currently return no rows. - Indeed: most reliable board, employer size, revenue and country on the row, countryIndeed picks the national site - Glassdoor: same countryIndeed setting as Indeed, browser-fetched, rows land in 1 to 3 minutes - LinkedIn: job_level for seniority, about 100 results per IP, residential proxy recommended - ZipRecruiter: US and Canada postings only - Naukri: experience_range, skills and company_rating for India pay curves - Bayt: Middle East coverage, search term only, no hoursOld or jobType filters ### What does a salary benchmarking workflow look like with scraped postings? A repeatable benchmark is a saved input plus a notebook. Start with searchTerms holding up to 5 related titles, so one run covers the whole job family and every row carries matched_search_term to show which query found it. Set location, pick the boards for your market, raise maxResults toward its cap of 100 per board per term, and set hoursOld to 168 so the sample reflects the current week. Deduplication across boards happens before rows are written, so a job listed on three boards contributes one observation. Export CSV or Excel from the Output tab, or pull the dataset by API into pandas. In the notebook, drop rows where salary_min is null, split by salary_currency, then compute the median, p25 and p75 of the midpoint by title, location and company_employees_label. A single run tops out at 4,000 rows (100 results x 8 boards x 5 terms), so paginate with offset or split by city when you need more. Schedule the saved input weekly and append each run keyed on id and scraped_at for a time series. - Up to 5 searchTerms per run, each row tagged with matched_search_term - maxResults up to 100 per board per term; hoursOld 168 keeps the sample to the current week - One observation per job: cross-board duplicates are merged before rows land - Export CSV, Excel or JSON, or fetch the rows over the Apify API - Compute median, p25 and p75 of the midpoint by title, location and employer size - Append weekly runs keyed on id and scraped_at for a trend line ### What filter and coverage limits affect a salary sample? Two board-side restrictions shape sample design. Indeed cannot combine hoursOld with jobType, isRemote or easyApply, and LinkedIn cannot combine hoursOld with easyApply, so a recency-filtered pull of remote-only Indeed roles takes two steps: pull by hours, then filter is_remote in your own code. Bayt ignores everything except the search term. Expect fewer rows than maxResults when a board has fewer matches or duplicates were removed; that is coverage working, not a failure. Speed also differs by board: Indeed and LinkedIn rows land 5 to 20 seconds after start, while Glassdoor, ZipRecruiter, Bayt and Naukri go through a real browser and take 1 to 3 minutes, streaming into the dataset as each board finishes. For the sample itself, treat a posted range as an advertised band, not a paid salary: it skews toward roles where pay transparency laws or competitive markets force disclosure. Reporting the share of rows with any salary alongside the median keeps the benchmark honest, and salary_source lets you publish a structured-only figure next to the broader one. - Indeed: hoursOld cannot pair with jobType, isRemote or easyApply; filter those columns after the pull - LinkedIn: hoursOld cannot pair with easyApply - Bayt: search term only, so apply recency and type filters in your notebook - Indeed and LinkedIn rows arrive in 5 to 20 seconds; browser-fetched boards take 1 to 3 minutes - Report the share of rows with published pay next to every median ### Steps 1. Open the Datapika job board scraper on Apify, enter up to 5 searchTerms for the job family, set location and countryIndeed, and select the boards that serve your market. 2. Set hoursOld to 168, raise maxResults to 100, and switch on enforceAnnualSalary if you want one annual column; leave it off to study hourly markets in their native unit. 3. Run it, watch Indeed and LinkedIn rows stream in within 5 to 20 seconds, then export CSV or Excel or pull the dataset by API. 4. In your notebook, drop rows without salary_min, split by salary_currency, filter on salary_source, and compute medians and percentiles by title, location and employer size; schedule the same input weekly for a trend line. ### FAQ Q: Does Glassdoor have an official API? A: Not one you can apply for today. Glassdoor closed public access to its developer API years ago (Zuplo on dev.to dates it to 2021, JobsPipe to 2022), and the help page that once held the API docs now describes Glassdoor Customer Insights, a subscription product for HR teams rather than a developer job-data API. Indeed retired its Publisher API too, and its remaining partner APIs push jobs in rather than read them out. Reading the public search pages is the practical route to salary rows. Q: How do I pull salary data from several job boards in one run? A: Run one Datapika sweep with your target titles in searchTerms and the boards for your market in sites. Each returned row carries salary_min, salary_max, salary_currency, salary_interval and salary_source where pay is published, plus employer size, revenue and rating for segmentation. Duplicates are merged across boards before rows are written, so the same posting never counts twice in your median. Q: Can I compare an hourly contract rate with an annual salary in the same dataset? A: Yes. Enable enforceAnnualSalary and every hourly figure is multiplied by 2,080 and every monthly figure by 12, so all three intervals land in one annual column. The original salary_interval still tells you which rows were converted. Currency is not converted, so split by salary_currency before comparing a London row with a New York row. Q: What does salary_source mean, and should I filter on it? A: salary_source is direct_data when the board exposed the pay range as structured data and description when the number was parsed from the posting text. For a published benchmark, compute one figure on direct_data rows only and a second on all rows with pay; if they diverge, spot check the parsed rows. Parsed hourly values matter most, since annual normalization multiplies any mistake by 2,080. Q: How much does a 1,000-row salary dataset cost? A: $5 per 1,000 rows, billed only for jobs that actually land in your dataset, so a run that finds nothing costs nothing beyond the platform start fee. A default run of 6 boards at 20 results each returns at most 120 rows. The scraper has 2,471 users, 381 of them active in the last 30 days, 27,458 runs and a 5.0 rating from 3 reviews (Apify, August 29, 2026). Q: Which countries can I benchmark salaries in? A: Indeed and Glassdoor take a countryIndeed code covering usa, uk, canada, australia, germany, france, india, singapore, uae and most other markets. LinkedIn is global, ZipRecruiter covers the US and Canada, Naukri covers India with experience ranges and skills, and Bayt covers the Middle East. Run one country per sweep so salary_currency stays uniform, then combine the exports with your own exchange rates. ## How to pull every job posted in the last 24 hours from six boards in one run Canonical: https://datapika.com/use-cases/fresh-jobs-last-24-hours Updated: 2026-08-29 Set hoursOld to 24 in the Datapika job board scraper and one run returns only postings published within the last day from LinkedIn, Indeed, Glassdoor, ZipRecruiter, Naukri and Bayt, deduplicated across boards and exported as JSON, CSV or Excel. Indeed and LinkedIn rows land in the dataset 5 to 20 seconds after start; browser-fetched boards follow within 1 to 3 minutes. Each job costs $0.005. The actor has 27,458 runs and 2,471 users on Apify with a 5.0 rating, and a daily schedule turns the same input into a standing fresh-jobs feed. ### How do I filter for jobs posted in the last 24 hours across every board? The hoursOld input takes a whole number of hours and is passed to each board's own recency filter, so 24 means the last day and 168 means the last week. Pair it with searchTerm, or up to 5 searchTerms for an OR search where every row carries matched_search_term, then add a location and a distance radius (default 50 miles). The default board list is LinkedIn, Indeed, Glassdoor, ZipRecruiter, Bayt and Naukri; Google Jobs stays selectable but currently returns nothing. maxResults caps each board at 1 to 100 rows per term, default 20, and a single run tops out at 4,000 rows across all 8 boards. Results stream as each board finishes: Indeed and LinkedIn usually within 5 to 20 seconds, Glassdoor, ZipRecruiter, Bayt and Naukri within 1 to 3 minutes because they run through a real browser. A posting that appears on three boards collapses to one row, so counts reflect unique jobs, and every row keeps date_posted, scraped_at, site and a direct apply URL when the board exposes one. - hoursOld: any whole number of hours, 24 for a daily sweep, 168 for weekly; the window is sent to every selected board - searchTerms: up to 5 queries per run, merged, each row tagged with matched_search_term - maxResults: 1 to 100 per board per term, default 20; use offset to page deeper - Streaming: Indeed and LinkedIn rows in 5 to 20 seconds, browser-fetched boards in 1 to 3 minutes - Dedup: one row per unique job even when it is listed on several boards - Export: JSON, CSV, Excel, XML or RSS from the dataset, or through the API ### Which filters conflict with a 24-hour window? Boards do not all accept a recency filter alongside other filters, and the combinations that fail return fewer rows rather than an error. On LinkedIn, hoursOld cannot be paired with easyApply in the same search. On Indeed, hoursOld cannot be combined with jobType, isRemote or easyApply, which is the most common reason a remote-jobs-posted-today run comes back thin. Bayt accepts only searchTerm, so hoursOld is ignored there and its rows need a date_posted check on your side. LinkedIn also rate-limits at roughly 100 results per IP, which is why residential proxy is the default; keep it on for any fresh sweep larger than a test. The practical pattern for a remote-only daily feed is two passes: run the boards with hoursOld set to 24 and no job-type or remote flags, then filter the export on is_remote and job_type, both returned on every row where the board provides them. Glassdoor and Indeed also need countryIndeed set to the market you care about; the default is usa. - LinkedIn: hoursOld and easyApply cannot be used in the same search - Indeed: hoursOld excludes jobType, isRemote and easyApply in the same search - Bayt: only searchTerm is honored, so filter Bayt rows on date_posted yourself - LinkedIn: about 100 results per IP; residential proxy stays on by default - Workaround: fetch with hoursOld alone, then filter on is_remote and job_type in the export - Set countryIndeed (usa, uk, canada, india, uae and others) so Indeed and Glassdoor search the right market ### How precise are posting timestamps on each board? The multi-board scraper writes date_posted as a calendar date, for example 2026-04-01, plus scraped_at as an ISO 8601 UTC timestamp, so recency inside a day is not visible in this dataset. Indeed dates postings the same day, which is why a 24-hour window is driven mostly by Indeed. ZipRecruiter loads new postings once a day in a batch, so the freshest rows it returns tend to be 12 to 48 hours old, and 48 hours is the smallest window that gives reliable two-board coverage. When you need the hour, use the companion Indeed and ZipRecruiter scraper: it reads the exact posted_at from the ZipRecruiter page, filters by that timestamp instead of the board's date filter, accepts hoursOld from 1 to 720 (default 48), and reports reposted_at separately when an employer bumps an old ad, so a refreshed listing never masquerades as new. It bills $0.0005 per record. Keep the multi-board actor for breadth and the two-board actor for hour-level precision and repost detection on the same search terms. - job-board-scraper: date_posted is day-level, scraped_at is a full UTC timestamp - Indeed: postings dated the same day, the anchor for any 24-hour sweep - ZipRecruiter: daily ingest batches, freshest jobs typically 12 to 48 hours old - indeed-ziprecruiter-scraper: hour-precise posted_at, hoursOld 1 to 720, default 48 - reposted_at flags employer bumps so reposts stay out of your new-today list - recencyMode last48h or latest10 on the two-board actor when volume matters more than age ### How do I schedule a daily fresh-jobs feed? Save the input once and attach an Apify schedule, for example 07:00 in your timezone, with hoursOld at 24 so consecutive runs tile the calendar with no gap; use 26 if you want a small overlap and dedupe downstream. Deduplication happens per run, across boards, so across days you should key your own store on id plus site and drop rows you have already seen. Attach a webhook to the run-finished event to push new rows into a Slack channel, Google Sheet, Airtable base or CRM through Zapier, Make or n8n, or read the dataset from the API with run-sync-get-dataset-items in a single call. Cost stays predictable because you pay only for rows delivered: cap the run with a maximum total charge in the Apify Console, keep maxResults near 20 while tuning, and remember that the Apify free plan's $5 monthly credit covers roughly 1,000 jobs. Runs default to 4 GB of memory, enough for the browser-fetched boards. For a weekly digest, switch hoursOld to 168 and raise maxResults. - Schedule in Apify Console with the saved input; hoursOld 24 tiles days end to end - Dedup is per run; across days, key on id and site in your own store - Webhook on run finish to Slack, Sheets, Airtable or a CRM via Zapier, Make or n8n - run-sync-get-dataset-items returns the rows in one API call - Set a max total charge per run; free plan credit covers about 1,000 jobs a month - Weekly variant: hoursOld 168 with a higher maxResults ### Who runs 24-hour job sweeps and what do they do with the rows? Job seekers and career coaches run one search every morning and apply from the direct apply URL while a posting is hours old and the applicant count is small. Recruiters watch for new openings at target accounts, using linkedinCompanyIds to restrict LinkedIn to named employers, and treat a burst of postings as a sourcing trigger. Sales teams read the same burst as a buying signal: five new SDR roles at one company is a reason to call, and the competitor-hiring guide covers scoring accounts on posting velocity. Niche job boards and aggregators refill their listings without maintaining six separate scrapers, and AI agents call the actor through the Apify MCP server with prompts such as remote React roles posted in the last 24 hours with salary ranges. Over the last 30 days, 381 users ran the actor on Apify. If you need boards beyond the six defaults, the full catalog at /scrape lists every board Datapika covers. - Job seekers: daily search, direct apply URL, first-wave applications - Recruiters: linkedinCompanyIds to watch named employers for new openings - Sales: posting bursts as buying signals; score accounts on new-role counts - Aggregators: refill listings from six boards with one scheduled run - AI agents: MCP tool at mcp.apify.com, natural-language prompts over fresh rows ### Steps 1. Open apify.com/openclawai/job-board-scraper, enter a searchTerm (or up to 5 searchTerms) and a location, and set hoursOld to 24. 2. Leave the six default boards selected, set countryIndeed for your market, and keep jobType, isRemote and easyApply off so Indeed and LinkedIn honor the window; filter those columns after export. 3. Run it, read Indeed and LinkedIn rows within 5 to 20 seconds, then export JSON, CSV or Excel or fetch the dataset through the API. 4. Attach a daily schedule and a webhook in Apify Console, dedupe across days on id and site, and add the Indeed and ZipRecruiter scraper when you need hour-precise posted_at and repost flags. ### FAQ Q: Is there one search that returns jobs posted in the last 24 hours from every job board? A: Run the Datapika job board scraper with hoursOld set to 24 and your search terms. One run sweeps LinkedIn, Indeed, Glassdoor, ZipRecruiter, Naukri and Bayt, merges duplicates into single rows, and streams results as each board finishes. Add up to 5 search terms for an OR search and raise maxResults to 100 per board when you need depth. Bayt ignores the window, so check its date_posted values before counting those rows as new. Q: Does Indeed have an official API for fresh job postings? A: No self-serve one. Indeed's Publisher job search API was retired, and the APIs documented at docs.indeed.com today cover job posting management, candidate handling and employer entities, with access provisioned by Indeed through its Partner Console. Job search is offered only as a hosted JavaScript plugin for publisher sites, not as a data endpoint. ZipRecruiter's API is likewise reserved for partner job boards and publishers that syndicate listings, not for on-demand job queries. The scraper reads the public search pages instead, so no partner agreement is required. Q: Why does a 24-hour run include jobs older than a day? A: Three causes. Bayt honors only searchTerm, so its rows are unfiltered. ZipRecruiter loads postings in a daily batch, so its freshest rows are typically 12 to 48 hours old. And employers bump old ads, which a board's own filter treats as new. Filter Bayt on date_posted, widen ZipRecruiter to 48 hours, and use the Indeed and ZipRecruiter scraper's reposted_at field when bumps must be excluded. Q: Can I get hour-precise posting times instead of a date? A: Not from the multi-board scraper, which returns date_posted as a calendar date and scraped_at as the run timestamp. For hour precision use the companion Indeed and ZipRecruiter scraper: it captures posted_at from the ZipRecruiter page, filters by that timestamp with hoursOld from 1 to 720, marks employer bumps as reposted_at, and accepts up to 10 keywords or a list of company names per run. Q: How much does a daily 24-hour sweep cost? A: You pay $0.005 per job delivered, $5 per 1,000, and a run that returns zero jobs costs nothing beyond the Apify platform's small start fee. A default run of six boards at 20 results each tops out at 120 rows, or $0.60, and most daily sweeps come in below that after deduplication. Put a spending cap on the run with the maximum total charge setting, and note that the Apify free plan includes $5 of monthly credit, about 1,000 jobs. Q: Why does LinkedIn return fewer fresh jobs than Indeed? A: LinkedIn rate-limits at roughly 100 results per IP, so larger sweeps need the residential proxy that is on by default, and it will not combine hoursOld with easyApply. Full descriptions and direct URLs also require linkedinFetchDescription, which adds requests and time. Indeed has no rate limit in the scraper and dates postings the same day, which makes it the higher-volume board for a strict 24-hour window. ## How to scrape Google Flights prices, routes, and CO2 data Canonical: https://datapika.com/scrape/google-flights Updated: 2026-08-29 There is no public Google Flights API: Google shut down QPX Express on April 10, 2018 and never replaced it. Datapika's Google Flights scraper on Apify fills that gap by running a route and date query against the live Google Flights results page and returning every itinerary as structured JSON. Google typically shows 30 to 80 itineraries per query, and each one costs $0.0015, so a single search runs roughly 5 to 12 cents and 1,000 itineraries cost $1.50. Each row carries the total price, airline IATA codes, stop count, per-segment airports and aircraft, and CO2 emissions in grams. ### Is there an official Google Flights API? No. Google's only public airfare feed was QPX Express, a product of its $700 million ITA Software acquisition. Google announced on November 1, 2017 that it would retire the API, citing low interest among travel partners, and switched it off on April 10, 2018. The enterprise version stayed available only to large contracted customers, so independent developers lost programmatic access to Google's fare data. That is the gap a Google Flights scraper fills: it reads the public results page for a route and date and converts what a traveler would see into machine-readable rows. Datapika runs on Apify with no API key on your side and no login to Google, and the store listing showed 3,696 runs as of August 2026. You pay per itinerary delivered rather than per month, which matters when you only need a few routes checked once a day. - QPX Express was retired on April 10, 2018 after a five-month notice period, per TechCrunch - Google cited low interest among travel partners as the reason for the shutdown - QPX Enterprise continued, but only for large contracted travel companies - Datapika reads the public Google Flights results page instead, no key or Google account needed - Billing is per itinerary returned, not per month or per API call ### What fields does a Google Flights scraper return per itinerary? Every dataset row is one itinerary for the route and date you asked for. At the top level you get the total price as a number, the list of airline names and their two-letter IATA codes, the stop count, total trip time in minutes, and ISO 8601 departure and arrival timestamps. The segments array then breaks the trip into legs, each with the origin and destination airport code and full name, leg departure and arrival times, leg duration, and the aircraft type as Google displays it, for example Airbus A320. Two carbon fields close the record: this itinerary's CO2 estimate in grams and the typical figure Google reports for the route, so you can flag greener options without extra math. A scraped_at UTC timestamp is stamped on every row, and the query inputs (airports, date, trip type, cabin, passenger counts) are echoed back so rows from different runs stay self-describing. - price (number), airline_codes (array of IATA codes), airline_names (array) - stops (integer, 0 = nonstop) and duration_min_total (integer minutes) - departure and arrival as ISO datetimes for the whole itinerary - segments: per-leg airport code and name, times, duration_min, plane_type - carbon_emission_g and carbon_typical_on_route_g in grams - scraped_at plus echoed inputs: from_airport, to_airport, date, return_date, trip, seat, passengers ### How much does it cost to scrape Google Flights? Datapika charges one event, itinerary scraped, at $0.0015, as listed on the Apify Store in August 2026. Nothing is charged for the query itself, so a route with no results costs nothing. Google Flights typically surfaces 30 to 80 itineraries per route and date, which puts a single search between about 4.5 and 12 cents. The max_results input defaults to 25 and caps at 200, so you control the ceiling on every run. At the default of 25 a search costs at most $0.0375, and 1,000 itineraries in any combination of routes cost $1.50. For a fare monitor that checks 20 routes twice a day at the default cap, that is 1,000 itineraries per day, or roughly $45 per month, with no base subscription. If you set a maximum total charge on the run in Apify, the actor stops pushing rows once that budget is reached, and the charge is posted once at the end for the rows actually delivered. - $0.0015 per itinerary returned, Apify Store, August 2026 - Typical single search: 30 to 80 itineraries, about 4.5 to 12 cents - max_results default 25 (at most $0.0375 per search), maximum 200 - 1,000 itineraries = $1.50; 10,000 = $15 - No charge for empty results or for the query itself ### Which routes, trip types, and cabins can you scrape? The input takes three-letter IATA airport codes for origin and destination, such as JFK to LAX or LHR to NRT, and a departure date in YYYY-MM-DD format. Leave the date blank and the actor searches 30 days from today. Add a return_date and set trip to round-trip for combined outbound and return pricing; multi-city is the third trip type. Cabin class is one of four values: economy, premium-economy, business, or first. Passenger counts follow Google's own model: 1 to 9 adults, up to 8 children aged 2 to 11, and infants under 2 split into in-seat and on-lap, up to 8 each, so family and group fares come back priced correctly. The language input accepts an ISO code such as en, es, fr, de or ja and defaults to en, with 30+ result languages supported, which changes airline and airport names in the output to match the locale. Runs default to 1,024 MB of memory and a 300 second timeout. - Origin and destination as 3-letter IATA codes (JFK, LAX, LHR, CDG, NRT) - Trip types: one-way, round-trip, multi-city - Cabins: economy, premium economy, business, first - Passengers: 1 to 9 adults, 0 to 8 children, infants in seat and on lap - Language: ISO code such as en, es, fr, de, ja, 30+ supported, default en - Date defaults to today plus 30 days when left empty ### How do developers and AI agents call the Google Flights scraper? Three entry points share the same actor. In the Apify Console you fill the form and download the dataset as JSON, CSV, or Excel. From code you call the Apify API with your token: POST the input object to the run-sync-get-dataset-items endpoint for openclawai~google-flights-scraper and the response body is the itinerary array, ready for a pandas DataFrame or a Postgres insert. For agents, the actor is exposed as an MCP tool through Apify's MCP server at mcp.apify.com, so a Claude or GPT based travel assistant can ask for JFK to LAX on a given date and receive typed rows without HTML parsing. Because the query inputs are echoed in every row, you can fan out dozens of route and date combinations in parallel runs and merge the datasets by from_airport, to_airport, and date. Schedule runs in Apify to build a daily fare history for a route. - Console: fill the form, run, export JSON, CSV, or Excel - REST: run-sync-get-dataset-items returns the itinerary array in one call - Python: pip install apify-client, then client.actor('openclawai/google-flights-scraper').call(run_input=...) - MCP: add the actor as a tool through Apify's MCP server at mcp.apify.com - Schedules and webhooks in Apify turn one query into a daily fare history ### Steps 1. Open https://apify.com/openclawai/google-flights-scraper and enter the origin and destination IATA codes (for example JFK and LAX), a departure date, and an optional return date. 2. Pick the trip type (one-way, round-trip, multi-city), a cabin class, and passenger counts, then set max_results between 1 and 200 to cap your spend at $0.0015 per itinerary. 3. Start the run and export the dataset as JSON, CSV, or Excel, or fetch it through the Apify API with the run-sync-get-dataset-items endpoint. 4. For AI agents, connect Apify's MCP server at https://mcp.apify.com and enable openclawai/google-flights-scraper as a tool so the assistant can request itineraries directly. ### FAQ Q: Does Google Flights have an official API, and what happened to it? A: No public API exists today. Google's QPX Express API, built on the ITA Software technology it bought for $700 million, was the only sanctioned airfare feed. Google announced the shutdown on November 1, 2017, citing low interest among travel partners, and closed it on April 10, 2018. Only the enterprise version survived for large contracted customers. Scraping the public Google Flights results page is now the practical route to this data for independent developers. Q: Is the flight data live or cached? A: Live. Every run performs a fresh Google Flights search at the moment you start it, and each row carries a scraped_at UTC timestamp so you know exactly when the fare was observed. There is no shared cache between customers. If you need price history, schedule the actor in Apify and append each run's dataset to your own store; fares change often enough that a snapshot from an hour ago can already be stale. Q: What currency is the price in, and does the output include booking links? A: The price field is the total Google displays for the itinerary, typically in USD, and it can be null when Google shows no fare for a result. The output does not include a booking or deep link, because Google Flights routes purchases through airline and agency sites that vary per itinerary. Use the airline codes, flight times, and airport codes to rebuild a search on the carrier's site or in your own booking flow. Q: How many itineraries does one search return, and can I get more? A: Google Flights typically surfaces 30 to 80 itineraries per route and date, and the actor returns up to the max_results you set, default 25, maximum 200. If Google shows fewer results than your cap, you are charged only for what is delivered. To cover more options, run separate searches for nearby dates or alternate airports and merge the datasets using the echoed from_airport, to_airport, and date fields. Q: Can I scrape round-trip and multi-city itineraries? A: Yes. Set trip to round-trip and supply a return_date to get combined outbound and return pricing, or set trip to multi-city. The segments array lists every leg with its own airports, times, duration, and aircraft type, and the top-level stops count tells you how many layovers the itinerary includes. All four cabin classes and full passenger mixes (adults, children, infants in seat, infants on lap) work with every trip type. Q: What are the run limits and failure modes? A: The default run uses 1,024 MB of memory with a 300 second timeout, which is ample for a single route and date. Google sometimes throttles repeated queries, so the actor retries each search up to 3 times with 2, 4 and 8 second backoff before failing the run with a clear status message; a failed run pushes no rows and bills nothing. Routes with no direct or connecting service return an empty dataset at no charge. ## Scrape Reddit posts and comments without the official API Canonical: https://datapika.com/scrape/reddit Updated: 2026-08-29 You can scrape Reddit without the API by running the Datapika Reddit actor on Apify. It reads public Reddit pages, so it needs no developer key, OAuth token, or login, and it never depended on the unauthenticated .json endpoints that began returning 403 in late May 2026. Pick one of six actions (scrape_subreddit, search_posts, search_comments, search_subreddits, fetch_post, or reddit_answers), set a limit up to 500, and export JSON or CSV. Pricing is per result: $2 per 1,000 posts, $3 per 1,000 with full comment trees, and the actor has 6,931 runs with a 4.9 rating (Apify Store, August 2026). ### Why did scraping Reddit with .json URLs stop working? Since late May 2026, appending .json to a Reddit URL returns 403 Forbidden for requests without a login (Crawlora, 2026; FetchLayer, 2026). Reddit had already closed self-service OAuth registration in November 2025, so the .json trick was the last free path left, and removing it broke most self-hosted Reddit scrapers, monitoring scripts, and RSS bridges in a single week. Some pipelines failed loudly with 403s; others kept returning HTTP 200 with empty JSON, so the break went unnoticed until dashboards ran dry. Logged-in sessions and the official OAuth Data API still work, but both require an approved app and a Reddit account. Datapika's Reddit actor never depended on those endpoints. It requests the same public pages a browser loads, with browser-grade TLS fingerprinting and 8 rotating user agents, and falls back to old.reddit.com when the main site refuses a request. The Apify Store page showed 100 percent of runs succeeding as of August 2026. If your pipeline started returning 403s or empty JSON this summer, the fix is to change the data source, not to add retries. - Cutoff: late May 2026; unauthenticated .json requests now return 403 Forbidden (Crawlora, 2026) - Self-service OAuth app registration closed in November 2025, so new keys need Reddit approval (FetchLayer, 2026) - Reddit has flagged RSS as the next surface it may close (Crawlora, 2026) - Datapika reads public HTML pages, not .json endpoints, so the cutoff did not change its data source - Store-reported success rate: 100 percent of runs succeeded, as of August 2026 - No Reddit account, developer key, or OAuth registration is involved at any step ### What can you scrape from Reddit without an API key? The actor bundles six actions behind one input schema, so you choose an action and pass the matching parameters. scrape_subreddit pulls a subreddit's feed sorted by hot, new, top, rising, or controversial, with a time filter from past hour to all time on top and controversial sorts and an includeComments switch. search_posts runs a keyword search across all of Reddit or inside one subreddit, sorted by relevance, new, top, or most comments. search_comments finds public comments matching a phrase, optionally scoped to one subreddit, which is the quickest way to catch brand mentions buried deep in threads. search_subreddits takes a topic and returns communities with mention counts and sample posts, useful when you do not yet know where a niche audience posts. fetch_post takes a permalink and returns that thread with its complete comment tree. The sixth action, reddit_answers, queries Reddit's own AI answer engine and returns a markdown answer with follow-up questions, source post IDs, and source subreddits. The actor's listing describes it as the only Apify actor exposing that feature. - scrape_subreddit: any public subreddit, sorted by hot, new, top, rising, or controversial, time filters from hour to all - search_posts: keyword search across Reddit or scoped to one subreddit, sorted by relevance, new, top, or comments - search_comments: public comments matching a phrase, with parent IDs for thread context - search_subreddits: communities ranked by mention count, each with sample posts - fetch_post: one permalink in, full thread with nested comments out - reddit_answers: Reddit AI Answers as markdown plus follow_ups and source lists ### What fields does each Reddit post and comment include? Every post row shares one flat schema whichever action produced it, so you can union a subreddit sweep and a keyword search without remapping columns. The core fields are post_id, permalink, subreddit_name, author_name, title, body, media URLs, num_comments, num_upvotes, and an ISO 8601 post_timestamp. When includeComments is on for scrape_subreddit, each row carries a comments array holding author_name, body, media, and parent_id for every comment, so you can rebuild the nested thread by joining parent_id to post and comment IDs. Comment trees are fetched by a configurable pool of 1 to 20 parallel workers (default 10), which keeps a 100-post run with comments in the range of minutes rather than hours. A post with 500 comments still means 500 page requests, so budget time accordingly. AI Answer rows use a different shape: markdown, follow_ups, source_posts, and source_subreddits. Whatever the action, output lands in an Apify dataset you can download as JSON, CSV, or Excel, or read through the dataset API. - Post identity: post_id, permalink, subreddit_name, author_name - Content: title, full body text, and an array of media URLs (images, video links) - Engagement: num_upvotes and num_comments at scrape time - Timing: post_timestamp in ISO 8601 UTC, for example 2026-04-01T12:00:00Z - Comments: author_name, body, media, parent_id per comment, nested to full depth - Export: JSON, CSV, Excel, or programmatic access through the Apify dataset API ### How much does it cost to scrape Reddit without the API? Pricing is per result delivered, with no subscription and no minimum, and you are billed only for rows actually returned. Posts without comments are $2 per 1,000, posts with full comment trees are $3 per 1,000, search results and discovered subreddits are $2 per 1,000, single post fetches are $3 per 1,000, and AI Answers are $10 per 1,000 queries (Apify Store, August 2026). A daily brand sweep pulling 500 search results and 50 threads with comments therefore costs about $1.15, before proxy bandwidth. For comparison, Trudax's Reddit Scraper Lite, one of the most-used Reddit actors on Apify, lists from $3.40 per 1,000 results with a 4.57 rating and 40,040 users as of August 29, 2026. Reddit's own commercial Data API is quoted at roughly $0.24 per 1,000 requests, but each request needs an approved OAuth app and is metered per call, not per row (Crawlora, 2026). Datapika's actor carries a 4.9 rating from 8 reviews and 6,931 runs as of August 2026. - Posts without comments: $2 per 1,000 (Apify Store, August 2026) - Posts with full comment trees: $3 per 1,000 - Search results and subreddits found: $2 per 1,000 each - Single post fetch: $3 per 1,000; Reddit AI Answers: $10 per 1,000 queries - Worked example: 500 search results plus 50 threads with comments is about $1.15 - Trudax Reddit Scraper Lite for comparison: from $3.40 per 1,000, 4.57 rating, as of August 29, 2026 ### How do you scrape Reddit at scale without getting blocked? Reddit rate-limits aggressively, and a naive script hits 429 errors within a few hundred requests. The actor handles this with exponential backoff and jitter, so a 429 delays the next request instead of ending the run. Every request carries a browser-grade TLS fingerprint and one of 8 rotating Windows, macOS, or Linux user agents. When www.reddit.com refuses a page, the actor retries against old.reddit.com, and as a last resort an archive layer can recover some deleted posts and comments. The default proxy configuration already uses Apify's residential proxy group, billed at $8 per GB through your Apify account; the README recommends keeping it for anything beyond test runs, since datacenter IPs draw far more Reddit blocks. Start with limit set to 10 to 50 to validate a query, then raise it toward the 500 cap. Turn includeComments off when you only need titles and scores, since that speeds the run and cuts the price from $3 to $2 per 1,000 posts. Private and restricted subreddits are out of scope because they require a login. - Exponential backoff with jitter on 429 responses, so runs slow down rather than fail - 8 rotating browser user agents across Windows, macOS, and Linux plus TLS fingerprinting - Fallback chain: www.reddit.com, then old.reddit.com, then an archive layer for deleted content - Residential proxies on by default at $8 per GB; the README cites millions of posts per day with them - Comment fetching uses 1 to 20 parallel workers with random jitter between requests - Test with limit 10 to 50, then scale to 500; includeComments off drops cost to $2 per 1,000 ### Steps 1. Open the Datapika Reddit actor at https://apify.com/openclawai/reddit-scraper, or attach it to an agent through https://mcp.apify.com/?tools=fetch-actor-details,openclawai/reddit-scraper, and choose an action: scrape_subreddit, search_posts, search_comments, search_subreddits, fetch_post, or reddit_answers. 2. Fill the inputs for that action, for example subreddit technology, sort top, timeFilter week, limit 100, and leave includeComments off for a first pass. Keep the default residential proxy group if you plan to go past a few hundred rows. 3. Start the run. Rows stream into the dataset as they are fetched, so check the first 10 to 50 to confirm the fields you need, then raise the limit or add includeComments for full threads. 4. Download the dataset as JSON, CSV, or Excel, or read it through the Apify dataset API. Save the input and schedule it daily to turn a one-off scrape into a running Reddit feed. ### FAQ Q: Does Reddit still have an official API, and what happened to the .json endpoints? A: Reddit's official Data API still exists but requires an approved OAuth app; commercial use is quoted near $0.24 per 1,000 requests (Crawlora, 2026). The unauthenticated .json endpoints, which were the free workaround, stopped serving anonymous requests in late May 2026 and now return 403 Forbidden without a login. Datapika does not use either path, so the change did not alter its data source. Q: Do I need a Reddit account, developer key, or my own proxies to run this? A: No account, developer key, or OAuth registration is needed, because the actor reads public Reddit pages rather than the API. You do not need your own proxies either: the default run configuration uses Apify's residential proxy group, billed at $8 per GB through your Apify account rather than through the per-result price. You can switch to datacenter proxies for small tests, at a higher risk of Reddit blocks. Q: Can it recover deleted Reddit posts or comments? A: Partially. The fallback chain ends at an archive layer that can return some deleted posts and comments when Reddit itself no longer serves them. Coverage depends on what that archive captured before deletion, so treat recovered content as best effort rather than complete. Recovered rows are billed at the same per-result rate as live rows. Q: Does it work on private, restricted, or quarantined subreddits? A: No. Private and restricted subreddits require a logged-in Reddit account with access, and the actor deliberately runs without authentication. Any public subreddit, public post, public comment, or public search result is in scope. If a subreddit went private after you scraped it, historical rows remain in your dataset but new runs against it will return nothing. Q: Can an AI agent call this Reddit scraper through MCP or the API? A: Yes. The actor is exposed over the Apify MCP server, so a Claude, GPT, or custom agent can call it as a tool at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/reddit-scraper, pass an action and query, and read the dataset back. The same input works through the Apify REST API for scheduled jobs. Billing stays per result, so an agent pulling 200 search results spends about $0.40. Q: Is scraping public Reddit data legal? A: The actor only accesses content Reddit serves to any anonymous visitor and does not bypass authentication or touch private data. Public availability does not remove your own obligations: Reddit's user agreement, copyright in user posts, and data-protection laws such as GDPR still apply to how you store and use the rows, especially author names. Review those rules for your jurisdiction before running at scale. ## How to scrape Airbnb listings, calendars, and reviews through one API Canonical: https://datapika.com/scrape/airbnb Updated: 2026-08-29 Datapika's Airbnb scraper API returns six record types from one Apify actor, and every one of them bills a flat $0.001 per delivered record: search listings, room details with a date-anchored stay quote, reviews, calendar months, host portfolio items, and Experiences (Apify Store pay-per-event billing, August 2026). You start from a pasted airbnb.com search URL or a latitude and longitude bounding box, add check-in and check-out dates to get nightly prices, and every row lands as flat JSON tagged with a record_type field. Empty and failed runs are never billed, 1,000 records of any type cost $1.00, and the actor has completed 488 runs. ### What Airbnb data can you scrape with this API? The actor covers the Airbnb surfaces most analytics and pricing teams need, and each one bills its own per-record event at the same flat $0.001 rate. Search modes return listing cards with nightly price, rating, review count, Superhost and Guest Favorite flags, coordinates, and images. Room details fetches the full listing page: description, amenities, host ID and name, a price quote for your dates, plus embedded reviews and calendar arrays. Reviews mode returns one row per review with rating, comment text, language, creation date, and reviewer name and location. Calendar mode returns month-by-month availability for each room URL, host portfolio mode returns every active listing for a numeric host ID, and Experiences mode searches Airbnb Experiences by place ID. All rows share the same flat shape, so a single dataset can hold listings and reviews side by side and be filtered on record_type in a spreadsheet, a BI tool, or an LLM pipeline. - Listing rows: room_id, room_url, title, room_type, person_capacity, bedrooms, beds, bathrooms, rating, review_count, is_superhost, is_guest_favorite, free_cancellation, price_per_night, price_total, latitude, longitude, images - Room rows add description, amenities, host_id, host_name, host_url, price_currency, and embedded reviews and calendar arrays - Review rows: room_url, review_id, rating, comment, language, created_at, reviewer_name, reviewer_location - Calendar rows carry room_url, room_id, and a per-room availability array billed per month scraped - Host portfolio rows reuse the listing shape with record_type set to host_listing and host_id attached - Every row includes record_type and a scraped_at ISO timestamp ### How do you scrape Airbnb listings by URL or map area? There are two ways to search for listings and both produce the same listing rows at $0.001 each. Search by URL is the fastest path: apply your filters on airbnb.com in a browser, copy the airbnb.com/s/ URL, and paste it as searchUrl. Since the May 2026 release the actor honors the location segment in the path, so a Paris URL returns Paris listings and multi-segment locations like New-York--NY--United-States work as well. Search by bounding box is built for map-driven analysis. You pass north-east and south-west latitude and longitude, an optional zoom level from 1 to 20, and filters for check-in and check-out, nightly price range, place type, amenity IDs, and free cancellation only. Both modes accept a currency code, a language code, and a maxResults cap that defaults to 100 and goes up to 5,000, which is the simplest way to keep a large sweep inside budget. - searchUrl: any airbnb.com/s/ URL with dates, guests, price, and map filters already applied - neLat, neLng, swLat, swLng plus zoom 1 to 20 (default 12) define the map area for bounding box search - placeType filters to Entire home/apt, Private room, Shared room, or Hotel room - priceMin and priceMax filter nightly rates in the currency you set (USD by default) - Bounding box search only returns prices when both checkIn and checkOut are set - maxResults defaults to 100 and caps at 5,000 listings per run ### How do you get Airbnb prices for specific dates? Airbnb prices depend on dates, guest count, and currency, so the actor anchors every quote to the stay you specify. In room details mode you pass up to 100 airbnb.com/rooms/ URLs, a checkIn and checkOut date in YYYY-MM-DD format, an adults count from 1 to 16, and a currency code. Each room row then carries price_per_night, price_total, and price_currency for exactly that stay, alongside rating, review count, host, amenities, images, and description. When a listing is not bookable for your dates, the row still ships with the price fields set to null instead of turning into an error row, so you keep the host, rating, amenity, and image data. Calendar mode is the complement for availability tracking: it returns one calendar record per room URL and bills $0.001 per month in that record, so a room whose calendar spans 12 months costs $0.012 and 100 such rooms cost $1.20. - room_details inputs: roomUrls (max 100 per run), checkIn, checkOut, adults (default 2, max 16), currency - Quote fields: price_per_night, price_total, price_currency, tied to the requested stay - Unavailable dates return null price fields rather than an error row - room_calendar returns one record per room URL, billed at $0.001 per calendar month - Reviews and calendar arrays embedded in room rows are best-effort, so a failing sub-fetch never drops the listing ### How much does it cost to scrape Airbnb? Pricing is per delivered record, and every record type bills the same flat $0.001 (Apify Store pay-per-event billing, August 2026). Search listings, host portfolio items, full room details with a stay quote, calendar months, reviews, and Experiences all cost $0.001 per row, so 1,000 records of any kind come to $1.00. Nothing is charged when a run returns no rows or fails outright. A typical market scan looks like this: 500 listings from a bounding box search cost $0.50, following up with room details on the top 100 costs $0.10, and pulling 1,000 reviews across those rooms costs another $1.00, for $1.60 in total. Each mode bills only its own event, so a run's total is simply the number of delivered records multiplied by $0.001. Keep runs scoped with maxResults and batch room URLs together rather than launching one run per listing. - listing-scraped: $0.001 per search result (search by URL or bounding box) - room-details-scraped: $0.001 per full room page with stay quote - review-scraped: $0.001 per review row - calendar-month-scraped: $0.001 per calendar month per room - host-listing-scraped: $0.001 per host portfolio item - experience-scraped: $0.001 per Experiences result ### How does the actor handle Airbnb blocking and partial failures? Airbnb fingerprints datacenter IP ranges aggressively, so the actor ships with residential proxy enabled by default and the README recommends leaving it on. Room details fetches retry up to 5 times with exponential backoff, and every retry rotates the proxy session so the exit IP changes and a soft block does not follow the request. Challenge and captcha pages are detected explicitly and trigger a fresh-IP retry instead of a misleading parse error. Batch runs degrade gracefully rather than failing as a whole. If some room URLs are blocked, the successful rows are delivered along with an info row that summarizes how many succeeded and how many to retry, and error rows carry a sanitized error_message, the mode, and the target input that failed. Per-listing sub-fetches for reviews, calendar, and host details are best-effort, so a hiccup on one endpoint returns an empty array and the rest of the row still ships. - Residential proxy on by default; Airbnb rate-limits datacenter IPs hard - Up to 5 retries per room with exponential backoff and a new proxy session each attempt - Challenge pages are detected and retried rather than surfaced as null-value errors - Info rows summarize partial batches; error rows include error_message, mode, and target - Every record has a scraped_at ISO timestamp for freshness checks downstream ### Steps 1. Open the Datapika Airbnb scraper on Apify (https://apify.com/openclawai/airbnb-scraper) and pick a mode: search_by_url, search_by_bbox, room_details, room_reviews, room_calendar, host_listings, or experience_search. 2. Fill the inputs for that mode: a pasted airbnb.com search URL or four bounding-box coordinates for search, up to 100 room URLs for details, reviews, or calendar, or a numeric hostId for portfolios. Add checkIn, checkOut, adults, and currency when you want stay quotes. 3. Set maxResults (default 100, max 5,000), leave the residential proxy on, and start the run. Rows arrive tagged with record_type and can be exported as JSON, CSV, or Excel or pulled through the Apify API. 4. For AI agents, connect the MCP endpoint at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/airbnb-scraper so an assistant can call the same modes as tools and pay per delivered record. ### FAQ Q: Does Airbnb have an official API I can use instead? A: Airbnb runs a partner API at developer.withairbnb.com, but access is limited to approved partners such as property management systems and channel managers, and Airbnb is not currently accepting new access requests from outside developers (Elfsight and Smoobu guides). There is no public key you can sign up for. Datapika's actor needs no Airbnb credentials: it reads public listing, review, and calendar pages and returns them as structured rows billed per record. Q: How does this compare to other Airbnb scrapers on Apify? A: The most used alternative on the Apify Store lists 16,383 users, a 4.52 rating, and a headline price from $1.00 per 1,000 listings, which matches the $1.00 per 1,000 listing rows here (apify.com/tri_angle/airbnb-scraper, checked August 2026). If you only need search cards at volume, the two are priced the same on that headline number. Datapika's case is breadth in one actor at one flat rate: date-anchored room quotes, reviews, per-month calendars, host portfolios, and Experiences all at $0.001 per record, with no charge for empty or failed runs. Q: How do I scrape every Airbnb listing in a city? A: Airbnb limits how many results a single search exposes, so one URL or one large bounding box will not return a whole city. Tile the area into several smaller bounding boxes, raise the zoom level toward 20 for dense districts, and run one search per tile with maxResults set high enough (up to 5,000). Deduplicate on room_id afterwards, since neighbouring tiles overlap at their edges. Q: Can I scrape Airbnb reviews in other languages? A: Yes. Reviews mode takes an ISO two-letter language code (en, es, fr, de, pt, ja, zh, and others) that sets the locale for the review request, and each review row carries its own language field, so you can filter by the language a review was written in. Pass up to 100 room URLs per run and use maxResults to cap reviews fetched per room. Each review row costs $0.001, so 1,000 reviews come to $1.00. Q: Can an AI agent call the Airbnb scraper through MCP? A: Yes. The actor is exposed as an MCP tool at mcp.apify.com with the openclawai/airbnb-scraper identifier, so agents built on Claude, ChatGPT, or any MCP-capable framework can run a bounding box search, request a room quote, or pull reviews as a tool call and pay per record. The flat row shape with record_type and scraped_at fields is designed to be parsed without a custom schema per mode. ## How to scrape Bilibili videos, comments, and creators from a BV id Canonical: https://datapika.com/scrape/bilibili Updated: 2026-08-29 To scrape Bilibili, paste a video URL (bilibili.com/video/BV...), a creator space URL, or a live room URL into Datapika's actor on Apify, pick a mode such as video_detail or video_comments, and start the run. Each video row costs $0.001 as of August 2026 on the Apify Store, so 1,000 Bilibili videos cost $1.00, and URLs that fail come back as unbilled error rows. Six of the actor's seven modes work on Bilibili, up to 5,000 items per URL, with play, like, reply, share, and favorite counts on every video_detail record and no login or API key required. ### Which Bilibili URLs and ids does the scraper accept? Bilibili identifies videos by BV ids, the 12-character strings that start with BV in every video URL, and the actor extracts them from any bilibili.com/video/ link. Short links of the form b23.tv/xxx are resolved to their canonical BV URL before scraping, so links copied from the Bilibili app work as pasted. Creators are addressed by their numeric uid through space.bilibili.com URLs, and live rooms by the numeric room id in live.bilibili.com URLs. You never have to set the platform field by hand. The default auto setting recognises bilibili.com and b23.tv hosts and routes the URL to the Bilibili pipeline, and trending mode with platform left on auto defaults to Bilibili's popular feed. An invalid or deleted BV id produces a 404 error row rather than a failed run, and error rows are not billed. - Video: https://www.bilibili.com/video/BV1GJ411x7h7 (BV id extracted automatically) - Short link: https://b23.tv/xxxx, resolved to the canonical BV URL first - Creator: https://space.bilibili.com/{uid}, used by user_posts and user_profile - Live room: https://live.bilibili.com/{room_id}, used by live_info - Popular feed: no URL at all, just mode trending with platform bilibili - Bulk: a urls array processed concurrently, with 500 links per run the recommended ceiling ### What fields come back for a Bilibili video? A video_detail run returns one flat JSON row per BV id in the same unified schema used for TikTok and Douyin, so downstream code does not branch on platform. Bilibili's stat object is mapped onto the shared engagement fields: view becomes play_count, like becomes like_count, reply becomes comment_count, share becomes share_count, and favorite becomes collect_count. Descriptive fields cover title, description, publish timestamp as ISO created_at, duration_sec, cover_url, and the creator's numeric mid plus display name. Hashtags are extracted from the description text. Two shared fields are always null on Bilibili, music_title and music_author, because the Bilibili video response carries no background audio object for the actor to map. The raw field carries the complete platform response, so anything Bilibili returns beyond the unified fields, such as multi-part page lists, stays available for your own parsing. The video_url_nowm field is filled by a separate play URL request when downloadVideos is on, and Bilibili rotates those tokens, so download promptly or fetch through your own client. - item_id is the BV id; url is the canonical bilibili.com/video/BV... link - play_count, like_count, comment_count, share_count, collect_count from Bilibili's stat object - title, description, created_at, duration_sec, cover_url, hashtags - author_id (numeric mid), author_name, author_username - video_url_nowm when a play URL is issued; tokens rotate, so some downloads fail - raw: the untouched platform response for fields the unified schema does not cover ### How do you scrape Bilibili comments and creator data? Set mode to video_comments and pass a video URL to pull the reply thread. Each comment row carries comment_id (Bilibili's rpid), reply_to_id (the parent rpid, null for top-level comments), comment_text, the commenter's mid and username, a like_count, and an ISO created_at. Pagination continues until maxItems is reached; the README example pulls the top 200 comments on one video with maxItems set to 200. For creators, user_posts walks a space URL page by page and returns one row per upload with title, description, cover, publish date, play count, and comment count. user_profile returns the creator's mid, display name, bio, avatar, and a verified flag derived from Bilibili's official account type. Bilibili's list endpoints return thinner records than the single-video endpoint. If you need like, share, and favorite counts for a creator's whole catalogue, collect the BV ids from user_posts, then run video_detail over them in a second bulk run. - video_comments: rpid, parent rpid, text, commenter mid, like count, timestamp - user_posts: title, description, cover_url, created_at, play_count, comment_count per upload - user_profile: mid, name, bio, avatar_url, verified (official account type) - live_info: room id, live_status, live_title, viewer_count, host name, cover - Comment and profile rows are billed as their own pay-per-event items, listed on the store page - user_likes is not available on Bilibili and returns a clear error row ### What can you do with Bilibili data for creator analytics and China market research? Bilibili reported 371 million monthly active users and 116.5 million average daily active users for the quarter ended June 30, 2026, in results released on August 27, 2026. It is the primary home of anime, gaming, and educational long-form video in China. Datapika gives you the numbers behind those audiences without a Chinese phone number or a Bilibili account. Analytics teams benchmark anime and gaming channels by running user_posts on a set of space URLs on a schedule, then tracking play_count and comment_count over time. Education creators and course sellers use video_comments to read how learners react to a series. Brands entering China pull the popular feed daily to see which formats and topics are gaining traction before committing to a campaign. Bilibili is also the better choice among the actor's platforms for time-sensitive monitoring, because Douyin's public creator feed can lag by around 6 days for some accounts while Bilibili's endpoints return current data. - Creator benchmarking: schedule user_posts across competing anime, gaming, or education channels - Audience research: video_comments on flagship uploads, then sentiment analysis on comment_text - Trend spotting: trending mode pulls Bilibili's popular feed as full video records - Live commerce and event monitoring: live_info for viewer_count and live_status on a room - Cross-platform comparison: the same schema lets you line Bilibili up against TikTok and Douyin rows ### How much does scraping Bilibili cost and how does it run at scale? Video rows cost $0.001 each as of August 2026 on the Apify Store, which works out to $1.00 per 1,000 videos from video_detail, user_posts, live_info, or trending. Comments and profiles are separate pay-per-event items with their own rates on the store's pricing tab. Error rows for deleted videos, bad BV ids, or anti-bot blocks are never billed, and you can cap a run with the ACTOR_MAX_TOTAL_CHARGE_USD setting so a large batch stops at your budget. The actor had recorded 1,340 runs on the store as of August 2026. maxItems defaults to 100 and accepts values from 1 to 5,000 per input URL. Bilibili requests run through an HTTP client pool rather than a browser, so they are fully concurrent, and a batch of 100 URLs typically finishes in 2 to 3 minutes. Keep runs at 500 URLs or fewer and split bigger jobs for parallelism. You can run the actor from the Apify console, call it over the Apify REST API, or expose it to an AI agent through the Apify MCP server. Residential proxy is on by default. - $0.001 per video, so 1,000 Bilibili videos cost $1.00 (Apify Store, August 2026) - Failed URLs return error rows and are not charged - maxItems from 1 to 5,000 per URL, default 100 - About 100 URLs in 2 to 3 minutes; recommended maximum 500 URLs per run - Console, REST API, and MCP access; export as JSON, CSV, or Excel from the dataset ### Steps 1. Open the actor at https://apify.com/openclawai/tiktok-douyin-bilibili-scraper (or add it to your agent via https://mcp.apify.com/?tools=fetch-actor-details,openclawai/tiktok-douyin-bilibili-scraper). 2. Choose a mode: video_detail or video_comments for a bilibili.com/video/BV... link, user_posts or user_profile for a space.bilibili.com/{uid} link, live_info for a live.bilibili.com room, or trending with platform set to bilibili and no URL. 3. Paste one url or a urls array, set maxItems (1 to 5,000), leave platform on auto, and start the run. 4. Read the dataset in the console or over the API; each row has platform bilibili, an item_type, the BV id or mid as item_id, and the unified stat fields, at $0.001 per delivered video row. ### FAQ Q: Does Bilibili have an official public API for third-party developers? A: The endpoints that third-party tools rely on are the same web endpoints the Bilibili site and app call, documented by the community rather than through a developer program, and they change without notice as Bilibili rotates its anti-bot defenses. Datapika wraps those endpoints, keeps up with changes, and turns any breakage into a structured error row instead of a failed run, so you do not have to maintain the integration yourself. Q: Can I download Bilibili videos through the scraper? A: Partly. When downloadVideos is on, the actor makes a separate play URL request for each BV id and fills video_url_nowm with the resulting stream link, preferring the highest-bandwidth DASH stream. Bilibili rotates the tokens on those links, so some downloads fail and the URLs expire quickly. Metadata and the rest of the row are always returned regardless. For reliable archiving, fetch the stream soon after the run or use the BV id with your own download client. Q: Why are follower counts empty on Bilibili profile rows? A: Bilibili serves follower, following, and upload counts from a separate statistics endpoint that the user_profile mode does not call at the moment, so follower_count, following_count, and video_count are null for Bilibili while name, bio, avatar_url, verified, and the numeric mid are populated. The raw field holds the full profile response. If you need a creator's upload volume, run user_posts on the same space URL and count the rows. Q: Does user_posts return full engagement stats for every upload? A: No. Bilibili's creator upload list returns thinner records than the single-video endpoint, so user_posts rows carry title, description, cover, publish date, play_count, and comment_count, while like_count, share_count, collect_count, and duration_sec are null. To get the full stat set for a whole channel, take the BV ids from the user_posts run and pass them as a urls array to video_detail. Both runs bill each video row at $0.001. Q: Do I need a Bilibili account, cookie, or Chinese proxy to scrape? A: No account is needed for public videos, comments, profiles, live rooms, or the popular feed, and the actor runs through Apify residential proxy by default with no extra setup. The optional cookie field exists for gated content and is not required for the Bilibili modes described here. Only public data is scraped, and you remain responsible for complying with Bilibili's terms and the privacy rules that apply to you. Q: Can I scrape Bilibili liked videos or search by keyword? A: Not in this actor. The user_likes mode is TikTok and Douyin only and returns an error row for Bilibili, and there is no keyword search mode for any platform. For discovery, use trending mode to pull Bilibili's popular feed, which returns full video records including creator mids you can then follow up with user_posts. If keyword search matters for your project, tell us through the actor's issues tab so we can prioritise it. ## How to scrape Douyin videos, comments, and creator profiles Canonical: https://datapika.com/scrape/douyin Updated: 2026-08-29 To scrape Douyin, paste a douyin.com video, creator, or live-room URL into Datapika's Douyin scraper on Apify, pick one of 7 modes, and read the results back as JSON through the REST API or MCP. Each delivered record costs $0.001 as of August 2026 (Apify Store), with no API key, no Douyin account, and up to 5,000 items per input URL. Every row uses the same field names as Datapika's TikTok and Bilibili output, so one pipeline covers all three platforms. Failed URLs return an error row and are never billed. ### Which Douyin URLs can you scrape, and what does each one return? Datapika reads three Douyin URL shapes and detects the platform from the domain, so you can leave the platform setting on auto for douyin.com and iesdouyin.com links. A video URL in the form douyin.com/video/ feeds video_detail and video_comments. A creator URL in the form douyin.com/user/ feeds user_posts, user_profile, and user_likes, and a live.douyin.com/ URL feeds live_info. The trending mode needs no URL, but it does need platform set to douyin explicitly. Left on auto, trending falls back to Bilibili's popular feed instead. Douyin's trending surface is a board of ranked search phrases with a hot_value score, not a video feed, so those rows carry item_type hot_search_keyword rather than video. Some newer short-form Douyin share links are not parsed yet. If a share link fails, open it once in a browser and paste the canonical douyin.com/video/ form instead. Every unsupported input still produces a structured error row instead of a failed run. - douyin.com/video/: single video detail, or its comment thread with replies - douyin.com/user/: latest posts, profile card, or liked videos for one creator - live.douyin.com/: live status, stream title, viewer count, host name, cover image - No URL plus platform=douyin: the hot-search keyword board with hot_value scores - iesdouyin.com links are recognized as Douyin automatically - Unparsed short links return an error row; paste the canonical video URL to retry ### Which fields does a Douyin video, comment, or profile row include? A Douyin video row carries the first line of the caption as title (capped at 200 characters) and the full caption as description, the aweme ID as item_id, a canonical douyin.com/video URL, the creator's sec_uid as author_id, nickname as author_name, and Douyin ID as author_username. It adds duration_sec, cover_url, music_title, music_author, created_at as an ISO 8601 timestamp, and a hashtags array whose extractor handles Chinese characters. Engagement comes back as play_count, like_count, comment_count, share_count, and collect_count, the last one being Douyin's save-to-favorites metric. A creator profile row adds follower_count, following_count, video_count, bio, avatar_url, and a verified flag that is true for both personal and enterprise verification badges. Comment rows include comment_id, reply_to_id for threading, comment_text, the commenter's identifiers, like_count, and created_at. Every row also includes the raw source response under raw and a scraped_at timestamp, so Douyin fields that are not mapped into the unified schema are still available for analysis. - Video: title, description, hashtags, music_title, music_author, duration_sec, cover_url, created_at - Engagement: play_count, like_count, comment_count, share_count, collect_count - Profile: follower_count, following_count, video_count, bio, avatar_url, verified - Comment: comment_id, reply_to_id, comment_text, like_count, created_at - Live: live_status, live_title, viewer_count, author_name, cover_url - raw and scraped_at on every row for unmapped platform fields and audit trails ### How much does scraping Douyin cost, and how do bulk runs work? Datapika charges $0.001 per delivered record as of August 2026 (Apify Store), which is $1.00 per 1,000 rows. Douyin URLs that fail because a video was deleted, a creator went private, or the platform blocked a request return an error row with a readable error_message and are not charged. There is no subscription and no API key fee on top of the per-record price. For bulk work, pass a list under urls instead of a single url. URLs are processed concurrently, and the maxItems setting (1 to 5,000, default 100) caps how many rows each paginated URL returns, so 20 creator URLs at maxItems 50 yields up to 1,000 video rows for about $1.00. A batch of 100 URLs typically finishes in 2 to 3 minutes, and runs of 500 URLs or fewer are the most reliable. You can also cap spend for a run with Apify's maximum total charge setting. When the cap is hit mid-batch, remaining URLs are written as error rows and the run ends cleanly instead of overspending. - $0.001 per delivered record (Apify Store, August 2026); the actor has 1,340 runs - Error rows are free: deleted, private, region-blocked, or anti-bot failures cost nothing - maxItems from 1 to 5,000 per input URL, default 100 - About 100 URLs per run in 2 to 3 minutes; keep runs under 500 URLs for best reliability - Spending cap: remaining URLs become error rows once the maximum total charge is reached ### What can you build with Douyin data for the Chinese market? Douyin is the domestic Chinese counterpart to TikTok, and its data answers questions that Western platforms cannot. Brands selling into China use user_posts across competitor accounts to watch launch cadence, caption language, and collect_count as a save-intent signal. Agencies running KOL discovery pull user_profile rows for a candidate list and rank by follower_count against average play_count from the same creators' recent posts. The hot-search board is a fast read on what mainland audiences are searching right now. Scheduling the trending mode every hour and diffing hot_value across snapshots surfaces rising phrases before they show up in English-language coverage. Comment threads in Chinese feed sentiment models directly, since comment_text and like_count come back per reply with threading intact. Because the schema is shared with Datapika's TikTok and Bilibili modes, the same brand can be compared across all three platforms in one dataset without remapping fields. That matters for teams that report on China and global audiences side by side. - Competitor launch tracking: user_posts on rival accounts, sorted by created_at and collect_count - KOL discovery: user_profile rows ranked by follower_count and recent play_count - Hot-search monitoring: hourly trending snapshots with hot_value deltas - Chinese-language sentiment: threaded comment_text with like_count per reply - Live commerce checks: live_info for viewer_count and stream titles during campaign windows - Cross-platform reports: identical field names across Douyin, TikTok, and Bilibili ### How do AI agents call the Douyin scraper through the API or MCP? The actor is exposed on Apify as a standard REST endpoint, so any HTTP client can start a run with a JSON body containing mode, url or urls, and maxItems, then read the dataset as JSON or CSV when the run finishes. Authentication is your Apify token only; no Douyin credentials or cookies are required for public content. For agents, the same actor is available as an MCP tool through Apify's MCP server at mcp.apify.com. An agent connects with the actor enabled as a tool, calls it with the Douyin URL and mode as arguments, and receives the unified rows back in the tool result. Billing stays per record, so an agent that asks for 30 comments pays $0.03. The cookie field is optional and only needed for gated Douyin content such as private collections or certain feeds. Residential proxy is the default and the recommended choice for Douyin. When Douyin flags an exit IP, the actor retries on a fresh residential IP up to 4 times before writing an error row. - REST: POST a JSON input to the actor endpoint, then fetch the dataset items - MCP: enable openclawai/tiktok-douyin-bilibili-scraper as a tool on mcp.apify.com - Inputs: mode, url or urls, maxItems, optional includeComments, cookie, proxyConfiguration - Public content needs no cookie; gated collections accept a pasted browser cookie - Flagged exit IPs are retried on a fresh residential IP up to 4 times per URL ### Steps 1. Open https://apify.com/openclawai/tiktok-douyin-bilibili-scraper and choose a mode: video_detail, video_comments, user_posts, user_profile, user_likes, live_info, or trending. 2. Paste a Douyin URL in canonical form (douyin.com/video/, douyin.com/user/, or live.douyin.com/), or a bulk list under urls, and set maxItems between 1 and 5,000. For trending, set platform to douyin and skip the URL. 3. Leave the platform on auto for URL modes, keep the default residential proxy, and add a cookie only if you need gated content. Start the run and watch rows land in the dataset at $0.001 each. 4. Export the dataset as JSON or CSV, or connect an agent to Apify's MCP server at https://mcp.apify.com with the actor enabled and call it as a tool. ### FAQ Q: Does Douyin have an official API, and can I use it instead? A: Douyin runs an Open Platform for registered developers, but it is built around apps that account owners authorize, not around reading arbitrary public content. You cannot pull competitor profiles, other creators' videos, or their comment threads from it, and onboarding runs through a Chinese-language developer registration. Datapika reads public Douyin pages directly, returns unified rows within a run, and needs no developer application or Douyin account. Q: Why do some Douyin no-watermark MP4 links return a 403? A: Douyin's CDN currently rate-limits direct no-watermark video URLs, so many of them answer 403 until the platform rotates them. For that reason the downloadVideos option is auto-disabled for Douyin, and video rows still return full metadata, cover images, and stats. If your own IP or cookie can fetch those files, set forceDownload to true to request the URLs anyway. TikTok and Bilibili links are not affected. Q: How fresh is Douyin user_posts data? A: For most creators, user_posts returns the latest public videos, but Douyin's public feed lags by roughly 6 days for some accounts, as documented in the actor's platform notes. If you need same-day detection of a specific new video, scrape it with video_detail as soon as you have its URL, or monitor the hot-search board, which updates continuously. Time-sensitive TikTok and Bilibili monitoring is not affected by this lag. Q: Do I need a Douyin account, cookie, or Chinese phone number? A: No. Videos, comments, creator profiles, live rooms, and the hot-search board are all public and are scraped through the built-in residential proxy without any login. The optional cookie field exists for gated content such as private collections or certain personalized feeds. If you do paste a cookie, it is stored as a secret input field and only applied to that run. Q: What does the Douyin trending mode actually return? A: Douyin's trending surface is a hot-search keyword board, not a video feed, so trending rows have item_type hot_search_keyword with the phrase in title and a numeric hot_value score for ranking. You must set platform to douyin; on auto, trending defaults to Bilibili. Set maxItems to control how many phrases you receive. To get videos for a trending phrase, scrape specific videos or creators that match it with video_detail or user_posts. Q: Can I scrape Douyin, TikTok, and Bilibili in the same run? A: Yes. Put URLs from all three platforms in the urls list and leave platform on auto; each URL is routed to the right platform and every row comes back with a platform field and the same column names. The mode must be valid for each platform, so avoid live_info or trending for TikTok URLs and user_likes for Bilibili URLs, which return error rows instead of data. ## How to scrape the Google Ads Transparency Center Canonical: https://datapika.com/scrape/google-ads-transparency Updated: 2026-08-29 To see every ad a competitor runs on Google, enter its brand name, domain, or AR advertiser ID into Datapika's Google Ads Transparency scraper on Apify. Each run returns up to 2,000 ads per advertiser as JSON rows: format (text, image, or video), first and last shown dates, days active, image or video URLs, the advertiser's legal entity, and per-country impression ranges where Google discloses them. Coverage spans 230+ countries and every Google surface: Search, YouTube, Display, Shopping, and Maps. Pricing is $0.001 per result with no run-start fee, so 100 ads with full reach data cost about $0.10. ### What does the Google Ads Transparency scraper return for each ad? Every ad becomes one JSON row, so you can filter, sort, and diff results without parsing HTML. The row carries the advertiser identity (AR advertiser ID, display name, and, with full details on, the registered legal entity), the creative identity (CR creative ID and a permalink that opens the live ad), and the format. Dates arrive as first_shown and last_shown, and the actor computes days_active from them so a 1,747-day banner stands out from a test launched last week. Media is included rather than just referenced. Image ads return the banner URL, text ads return a rendered screenshot, and video ads return a renderable preview plus a direct video link whenever Google exposes one. Keep fetchAdDetails on (the default) to add variation_count, total_reach_low and total_reach_high, and a region_stats array with one entry per country the ad ran in. - advertiser_id, advertiser_name, legal_name: verified identity, for example AR0293… and Booking.com B.V. - creative_id and ad_url: a permalink into the Transparency Center that opens the live ad - format: TEXT, IMAGE, or VIDEO, filterable at input time so unwanted formats are never billed - first_shown, last_shown, days_active: run window and longevity for each creative - image_url, preview_url, video_url: real media, including rendered screenshots for text ads - variation_count, total_reach_low, total_reach_high, region_stats: A/B variants and impression ranges ### How do I look up an advertiser by brand, domain, or advertiser ID? The queries input takes brand names, website domains, AR advertiser IDs, and full Transparency Center URLs mixed in one list, and the actor detects the type of each line. Domains are the most reliable input because they resolve directly to the one verified advertiser that owns the site. Advertiser IDs and URLs also resolve to exactly one advertiser. Brand-name queries go through Google's suggestion search, which can return up to 20 advertisers sharing a name. The advertisersPerQuery setting (default 3, maximum 20) caps how many you keep, so a search for a common word does not pull in unrelated small businesses. Each matched advertiser is then paged newest-first in batches of 40 ads until maxAdsPerAdvertiser is reached, with a ceiling of 2,000 ads per advertiser per run. Results cover every Google surface, including Search, YouTube, Display, Shopping, and Maps. - Mix inputs freely: nike, booking.com, and AR02934798844673654785 can sit in the same queries list - Domains, IDs, and URLs resolve to one verified advertiser; brand names may match up to 20 - advertisersPerQuery defaults to 3 and caps at 20 for name searches - maxAdsPerAdvertiser defaults to 50 and caps at 2,000, fetched newest-first in pages of 40 - formats keeps only TEXT, IMAGE, or VIDEO ads; region limits results to one ISO country code ### Which countries does a competitor advertise in, and for how long? The region_stats array is the part most Transparency Center scrapers skip. For every country an ad ran in, you get the ISO code and name, a reach_low and reach_high impression range, that country's own first_shown and last_shown dates, and a surfaces list that splits the reach by surface code. Summed across ads, this shows where an advertiser concentrates budget and which creatives it scales into new markets. Google only publishes impression ranges where regulation requires it, mainly for ads shown in the EU, so reach fields can come back null for ads shown only in the US. Dates and days_active are populated regardless of region. Coverage spans 230+ countries, and the region input filters a run to a single country, which is also the practical way to segment advertisers whose inventory exceeds the 2,000 ads per run cap. - region_stats per country: code, name, reach_low, reach_high, first_shown, last_shown, surfaces - total_reach_low and total_reach_high give the ad-wide impression range where disclosed - days_active separates evergreen winners running for hundreds of days from creatives launched this week - Reach ranges are populated mainly for EU-shown ads; elsewhere they return null - Set region to DE, FR, or any 2-letter ISO code to pull one market at a time ### How much does it cost to scrape the Google Ads Transparency Center? Datapika charges one event: $0.001 per result, where a result is one dataset row, either an ad or an advertiser record (Apify Store, August 2026). Reach data, creatives, dates, and the legal entity are included in that price, there is no run-start fee, and ads dropped by the formats filter are never billed. Hint and error rows are free. Worked examples from the actor's documentation: one advertiser with 100 ads and full reach data is 101 results, or $0.10. Screening 500 domains with fetchAds set to false produces 500 advertiser records for $0.50. For comparison, as of August 2026 the automation-lab Google Ads Scraper on Apify lists $0.001 per ad plus $0.005 per run start, and the whoareyouanas Google Ads Transparency Scraper starts at $5.00 per 1,000 ads with a full mode near $0.015 per ad. The per-ad rate here matches the cheapest of those while including per-country reach. - $0.001 per result flat, no actor-start fee, no separate detail upcharge - 100 ads plus 1 advertiser record: $0.10 - 500-domain screening run with fetchAds off: $0.50 - Ads filtered out by formats are not charged; hints and errors are free - Competitor listing prices verified August 2026: $0.001 per ad plus $0.005 per run start; from $5.00 per 1,000 ads ### How do I run it on a schedule or from an AI agent without getting blocked? The Transparency Center rate-limits a single IP after roughly 25 requests, which is why home-grown scripts stall partway through a large advertiser. The actor defaults to Apify residential proxies and rotates IPs proactively during the run, retrying failed pages automatically, so a 2,000-ad pull completes without babysitting. No Google login, cookies, or API keys are involved because the source is the public Transparency Center. For monitoring, schedule a daily run on Apify with your competitor domains in queries and diff the dataset against yesterday's: new creative_ids are fresh tests, and rising days_active marks the winners. Agents can call the same actor through the Apify API or the MCP endpoint, passing a domain and reading back structured rows, which turns "list every ad this company is running" into a single tool call. Exports go to JSON, CSV, Excel, Google Sheets, Make, Zapier, or LangChain like any other Apify actor. - Residential proxy on by default; the platform throttles a single IP after about 25 requests - Automatic IP rotation and retries; no login, cookies, or API keys required - Schedule daily runs and diff creative_id sets to catch new tests the day they launch - MCP endpoint: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/google-ads-transparency-scraper - Exports: JSON, CSV, Excel, Google Sheets, Make, Zapier, LangChain ### Steps 1. Open https://apify.com/openclawai/google-ads-transparency-scraper and add your targets to queries: a brand name such as nike, a domain such as booking.com, or an AR advertiser ID. 2. Set maxAdsPerAdvertiser (default 50, up to 2,000), keep fetchAdDetails on for reach and legal entity data, and optionally restrict formats to VIDEO or region to a single ISO code such as DE. 3. Run it and read the dataset: one row per ad with format, first_shown, last_shown, days_active, media URLs, and region_stats. Export as JSON or CSV, or pull rows through the Apify API. 4. For agents, connect the MCP endpoint at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/google-ads-transparency-scraper and call the actor with a domain to get the same rows back as structured output. ### FAQ Q: Does Google offer an official Ads Transparency Center API? A: No. Google publishes no API for the general Ads Transparency Center. The only official feed is the google_political_ads public dataset on BigQuery, which covers political advertising only and reports spend and impression ranges plus creative metadata rather than the ads themselves (Google's political ads transparency announcement and adlibrary.com's overview, checked August 2026). Commercial advertisers such as Nike or Booking.com are absent from it. Datapika reads the public Transparency Center pages instead and returns the creatives, dates, and reach ranges a person would see in the browser. Q: Why are total_reach and region_stats null for some ads? A: Google discloses impression ranges only where regulation requires it, mainly for ads shown in the EU. An ad that ran only in the US or another non-EU market still returns its format, first_shown, last_shown, days_active, and media URLs, but reach_low and reach_high stay null. If reach is what you need, set region to an EU country code such as DE or FR so the run focuses on ads where the ranges are published. Q: Why did a brand-name search return advertisers I did not expect? A: Brand-name queries use Google's suggestion search, which can list up to 20 advertisers sharing a name, including small businesses with similar names. Keep advertisersPerQuery low (the default is 3) to stay focused, or switch to the domain form of the query, for example nike.com instead of nike. Domains, AR advertiser IDs, and Transparency Center URLs each resolve to exactly one verified advertiser, so they avoid the ambiguity entirely. Q: Can I pull every ad from an advertiser with tens of thousands of creatives? A: A single run returns up to 2,000 ads per advertiser, fetched newest-first in pages of 40. For larger inventories, split the work with the region filter and run one country per job, or run repeatedly and dedupe on creative_id. Advertiser rows also carry declared_ad_count, Google's own total for that advertiser, so you can see how much of the inventory a run captured. Each ad is billed once at $0.001. Q: Why does a video ad have a preview URL but no video URL? A: Google serves some video creatives only as renderable previews rather than as direct media files. When a direct video link is exposed, the actor includes it in video_url; when it is not, you still get preview_url, which renders the ad, plus the format, dates, days_active, and reach fields. Text ads are handled differently: they return a rendered screenshot in image_url so the copy is captured as an image. Q: Is scraping the Google Ads Transparency Center legal? A: The actor only reads data Google publishes for transparency purposes, available to anyone without a login, and it does not touch private Google Ads accounts or cookies. That said, how you use ad data depends on your jurisdiction and purpose, so review your own compliance requirements before storing or redistributing creatives. The output includes the advertiser's legal entity name, which helps when you need to attribute an ad correctly in a report. ## How to scrape Google Maps places, reviews, and contact data Canonical: https://datapika.com/scrape/google-maps Updated: 2026-08-29 Datapika's Google Maps scraper API returns structured place records, business name, address, phone, website, rating, review count, opening hours, coordinates, and optionally the top 10 reviews, from any Maps search query, with no Google API key. Pricing is a flat $0.003 per place with no volume tiers, so 1,000 places cost $3.00 and 5,000 cost $15.00. Google Maps caps a single search at roughly 120 places, so larger datasets come from running several narrower queries. The actor runs on Apify via REST API, CLI, or MCP server. ### What data does the Google Maps scraper API return? Every place in the dataset is a flat JSON record built from the Google Maps listing, not a raw HTML page. The core fields are title, category, address, phone, website, review_rating, review_count, status, description, and price_range. Location comes as latitude, longitude, plus_code, and an IANA timezone, and every row keeps a direct link back to the place on Google Maps plus Google's numeric cid identifier. Two fields are lists rather than scalars. open_hours is keyed by weekday, and images holds up to 5 photo URLs per place. reviews is null unless you enable review scraping, in which case it carries the top 10 reviews for that listing. Each record also stores search_query and an ISO scraped_at timestamp, so you can merge several runs into one table and still know which query produced which row and when. The dataset exports as JSON, CSV, or Excel, or you can read it item by item through the Apify API. - Contact fields: phone, website, and full street address for lead lists - Reputation fields: review_rating on a 1.0 to 5.0 scale plus total review_count - Location fields: latitude, longitude, plus_code, and timezone for mapping - Hours and status: open_hours per weekday and an Open / Closed / Temporarily closed status - Media: up to 5 image URLs per place, plus the top 10 reviews when enabled - Provenance: search_query and scraped_at on every row ### How much does it cost to scrape Google Maps places? Pricing is pay per result at a flat $0.003 per place, taken from the actor's live Apify billing as of August 2026. There are no volume tiers: the first place and the ten thousandth place cost the same. The only other event is an actor start charge of $0.00005 per run, which is small enough to ignore in any budget. You are billed only for places actually written to the dataset, and no Google Cloud account or API key is involved. The math is easy to check. A quick 50-place test costs $0.15, a full single query of about 120 places costs $0.36, and 1,000 places cost $3.00. At 5,000 places the run costs $15.00, and 10,000 places cost $30.00, which is the same $0.003 per place at every scale. The Apify Store lists the actor at $0.003 per place as its headline event price and shows 559 completed runs as of August 2026. The billing lists no separate per-review price; the input notes flag review scraping as slower, so budget extra run time rather than extra spend per place. - Flat rate: $0.003 per place at any volume (live Apify billing, August 2026) - Actor start event: $0.00005 per run, negligible in practice - No volume tiers, so the per-place cost never changes with run size - Worked examples: 1,000 places = $3.00, 5,000 = $15.00, 10,000 = $30.00 - You pay only for places delivered to the dataset, per the pay-per-result model ### Why does one Google Maps search stop at about 120 places? Google Maps itself renders only about 120 results for any one search no matter how far you scroll, and the actor's input notes flag the same ceiling. Setting maxResults to 500 on a single broad query like restaurants in New York will not return 500 places, it will return roughly 120 and stop. This is a property of Google Maps, not a quota on the actor. The way around it is narrower queries. Split by neighborhood (restaurants in Williamsburg Brooklyn), by category (vegan restaurants in New York, pizza in New York), or by both, and run one query per slice. Ten neighborhood queries at up to 120 places each yield close to 1,000 places, with each run billed at the same flat $0.003 per place. Deduplicate on the cid or link field when merging slices, because a business near a neighborhood boundary can appear in two searches. Concurrency defaults to 5 parallel place fetches and the input caps it at 10, the level where the actor's notes say Google rate limiting begins. - Roughly 120 places per query is a Google Maps display limit, not an actor limit - maxResults accepts 1 to 500, default 20, but a single query rarely exceeds about 120 - Split by neighborhood, district, postcode, or sub-category to multiply coverage - Merge runs on the cid or link field to drop duplicates across overlapping areas - Keep concurrency at the default 5; the input caps it at 10 to avoid rate limiting ### How do you scrape Google Maps reviews with the actor? Reviews are off by default because they add work per place. Set scrapeReviews to true and the actor opens each listing's review panel and captures the top 10 reviews, each with reviewer name, star rating, date, and review text. The input notes put the overhead at about 3 seconds per place, so 100 places with reviews add roughly 5 minutes to a run. The reviews field is an array nested inside the place record, so a 20-place coffee shop query returns 20 rows and up to 200 review objects. There is no separate per-review charge in the actor's billing, the flat $0.003 place price covers them. Ten reviews per place is enough to sample sentiment, catch recent complaints, or feed a summarization step in an agent pipeline. It is not a full review history, so for a business showing 843 reviews you get the 10 that Google surfaces first, not all 843. - Enable with scrapeReviews: true in the input JSON - Each review carries reviewer name, rating, date, and text - Up to 10 reviews per place, nested in the reviews array - Adds about 3 seconds per place, per the actor's input notes - Pair with review_count and review_rating for a complete reputation snapshot ### How do AI agents and scripts call the Google Maps scraper? The actor is built to be called by code and by agents, not only from the Apify console. Over REST you POST the input JSON to the actor's run-sync-get-dataset-items endpoint with your Apify token and receive the place records in the same response, which fits a serverless function or a cron job. The Apify CLI and the JavaScript and Python client libraries wrap the same call. For AI agents, the Apify MCP server exposes this actor as a tool. Point Claude, Cursor, or any MCP-compatible client at the MCP URL with openclawai/google-maps-scraper in the tools list, and the agent can read the input schema, run a query, and consume the dataset without custom glue code. The input is small enough for an agent to construct reliably. query is the only required field, while maxResults (default 20, maximum 500), concurrency, scrapeReviews, and proxyConfiguration are optional. Residential proxies are prefilled in the default input because Google's bot detection is aggressive. - REST: one POST to run-sync-get-dataset-items returns the full dataset - MCP: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/google-maps-scraper - Only query is required; maxResults defaults to 20 and caps at 500 - concurrency defaults to 5 and the input caps it at 10, where rate limits start - Residential proxy configuration is prefilled by default ### Steps 1. Open https://apify.com/openclawai/google-maps-scraper and sign in to Apify, or add the actor to your agent with https://mcp.apify.com/?tools=fetch-actor-details,openclawai/google-maps-scraper. 2. Set query to a specific search such as dentists in London, keep maxResults at or below about 120 per query, and turn on scrapeReviews if you need review text. 3. Run from the console, the REST API, the CLI, or an agent tool call; places are written to the dataset as each one completes. 4. Export JSON or CSV, deduplicate on the cid or link field when merging several neighborhood or category queries, and schedule the run if you need recurring snapshots. ### FAQ Q: Does Google Maps have an official API, and why not use it directly? A: Yes. The Google Places API is the official route, but it needs a Google Cloud billing account and is priced per request by SKU. Google's pricing page (updated 25 August 2026) lists Text Search Pro at $32 per 1,000 requests and Place Details Pro at $17 per 1,000 after 5,000 free monthly calls each. Getting contact data for 1,000 places means a search call plus a details call per place. Datapika returns the same 1,000 records for $3.00 with no key. Q: How many Google Maps results can I scrape in one run? A: Around 120 places per search query, because that is roughly how many results Google Maps renders for any single search. The maxResults input accepts up to 500, but a broad query will still stop near 120. To collect thousands of places, run one query per neighborhood, postcode, or sub-category and merge the datasets, dropping duplicates on the cid or link field. Ten focused queries can approach 1,000 unique places. Q: Can I scrape Google Maps reviews with this scraper API? A: Yes, by setting scrapeReviews to true. Each place then includes an array of up to 10 reviews with reviewer name, star rating, date, and full review text, nested inside the place record. It adds about 3 seconds per place and there is no extra per-review charge in the actor's billing. It is a top-10 sample, not the complete review history, so use review_count for the total volume. Q: Do I need proxies to scrape Google Maps? A: For anything beyond a small test, yes. Google Maps uses aggressive bot detection, and the actor's README recommends Apify residential proxies at scale. The default input already includes a residential proxy configuration, so you do not have to set anything up, and country-specific proxy sessions can be used to get localized results. Keep concurrency at the default of 5; the input caps it at 10, the point where the actor's notes say rate limiting starts. Q: What are the limitations of the Datapika Google Maps scraper? A: Three main ones. Each query tops out at roughly 120 places, so wide coverage requires many narrow queries. Reviews are capped at the top 10 per place. The output covers business-level data (name, address, phone, website, rating, hours, coordinates, images) but does not include email addresses or reviewer profile details, so an email-enrichment step is needed for outreach lists. Fields like open_hours or price_range are empty when the listing itself lacks them. ## How to scrape TikTok videos, comments, and profiles Canonical: https://datapika.com/scrape/tiktok Updated: 2026-08-29 Datapika scrapes TikTok videos for $0.001 each, which is $1 per 1,000 videos, with no TikTok API key or developer account (Apify Store, August 2026). Paste a video URL, a profile URL, or a vm.tiktok.com share link, pick one of the five TikTok modes, and get JSON rows with play, like, comment, and share counts, hashtags, music credits, and a no-watermark MP4 link. The same actor handles Douyin and Bilibili URLs in the same run, returns up to 5,000 items per URL, and is callable through the Apify REST API or an MCP endpoint. ### Which TikTok URLs and modes does the scraper accept? On TikTok the actor runs five of its seven modes: video_detail, user_posts, user_profile, video_comments, and user_likes. Each mode expects a specific URL shape, and the platform field can stay on auto because the actor detects tiktok.com, douyin.com, and bilibili.com hosts by itself. Video modes take the canonical form www.tiktok.com/@handle/video/1234567890, and profile modes take www.tiktok.com/@handle. Short links copied from the mobile app, vm.tiktok.com and vt.tiktok.com, are followed to their canonical video URL before scraping, so you can paste share links straight from a phone. There is no keyword or hashtag search input. Hashtags arrive as an array parsed from every caption, so the practical pattern is to scrape creators or videos you already know and filter on that array. The trending mode exists, but it serves Douyin hot search and Bilibili popular feeds, not TikTok. - Single url or a urls array; the two are merged if both are set, and 500 URLs per run is the recommended ceiling - maxItems caps paginated modes (user_posts, video_comments, user_likes) at 1 to 5,000 per input URL, default 100 - vm.tiktok.com and vt.tiktok.com share links resolve automatically to www.tiktok.com/@handle/video/ - Unsupported combinations such as live_info or trending on TikTok return a structured error row instead of failing the run - includeComments: true attaches the first page of comments to a video_detail run without a second call ### What fields come back for each TikTok video, profile, and comment? Every row shares one schema across the three platforms, which means a TikTok video and a Douyin video have identical field names and can land in the same table. The platform and item_type fields tell you what each row is: video, user, comment, live, trending, or error. Video rows carry title and description, author_id, author_name and author_username, duration_sec, cover_url, and video_url_nowm for the no-watermark MP4. Engagement comes as play_count, like_count, comment_count, share_count, and collect_count, with created_at as an ISO timestamp and music_title plus music_author for the sound. Profile rows add follower_count, following_count, video_count, bio, avatar_url, and a verified flag. Comment rows carry comment_id, reply_to_id, and comment_text so threads can be rebuilt. Each row also includes raw, the full source response, and scraped_at for freshness tracking. - Engagement: play_count, like_count, comment_count, share_count, collect_count - Creator: author_username, follower_count, following_count, video_count, verified, bio - Assets: cover_url, video_url_nowm (no-watermark MP4, expires within hours), duration_sec - Sound: music_title and music_author for tracking which audio is spreading - Threads: comment_id and reply_to_id link replies to their parent comment - raw keeps the complete source payload for fields the unified schema does not map ### How much does it cost to scrape TikTok compared with other scrapers? Datapika bills per delivered row. A video row costs $0.001, so 1,000 TikTok videos cost $1.00 and a full 5,000-item user_posts pull costs $5.00 (Apify Store, August 2026). Comments and profiles are separate pay-per-event items listed on the store page. There is no subscription, no API key fee, and no charge for URLs that come back as error rows. The category leader on Apify, clockworks/tiktok-scraper, lists $1.70 per 1,000 results, 245,964 users, and a 4.79 rating as of August 2026 (apify.com/clockworks/tiktok-scraper). It wins on adoption and on input types, since it accepts hashtags and search queries, which Datapika does not. Datapika wins on price per video and on covering Douyin and Bilibili in the same run. To cap spend, set ACTOR_MAX_TOTAL_CHARGE_USD on the run. When the cap is reached mid-batch, remaining URLs are skipped and written as error rows so the dataset still explains what happened. - $0.001 per video row, the same price on TikTok, Douyin, and Bilibili - 1,000 videos for $1.00; 5,000 videos, the maxItems ceiling for one URL, for $5.00 - Error rows (deleted, private, region-blocked, unsupported mode) are never billed - Comparison: $1.70 per 1,000 results on clockworks/tiktok-scraper, verified August 2026 - Spending cap via ACTOR_MAX_TOTAL_CHARGE_USD; the run stops charging at the limit ### How do marketing and research teams use TikTok data from this actor? The Western marketing case is creator work: vetting influencers, measuring sponsored posts, and spotting sounds before they peak. A user_profile call returns follower_count, video_count, and verified for a shortlist of creators, and a user_posts call on the same handles returns the last 100 posts with per-video engagement, enough to compute average views per post before signing a contract. For campaign measurement, run video_detail on the sponsored video URLs daily and store play_count, like_count, share_count, and collect_count by scraped_at. The deltas give a growth curve per post without asking the creator for screenshots. Sound sourcing uses music_title and music_author across a set of trend accounts, and comment mining pairs video_comments with a sentiment model. Because Douyin rows share the schema, a brand active in both markets can watch its Chinese creator partners in the same dataset. - Influencer vetting: follower_count, video_count, and average play_count from the last 100 posts - Sponsored post tracking: daily video_detail snapshots keyed by scraped_at - Sound and hashtag discovery: aggregate music_title and the hashtags array across trend accounts - Comment mining: video_comments with replies for sentiment and UGC sourcing (requires a TikTok cookie) - Cross-market benchmarking: TikTok and Douyin rows in one table with identical field names - AI and ML datasets: captions, stats, and cover images at $1 per 1,000 videos ### How do AI agents call the TikTok scraper through MCP? The actor is exposed through Apify's MCP server, so an agent can discover its input schema and run it without custom glue code. Point the client at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/tiktok-douyin-bilibili-scraper, call fetch-actor-details once to read the mode and URL fields, then call the actor with mode, url or urls, and maxItems. For plain HTTP, the Apify REST endpoint run-sync-get-dataset-items returns the dataset as a JSON array in one request. Each element is one row with item_type set, so an agent can branch on video, user, comment, or error without parsing HTML. Bulk runs process URLs concurrently, and 100 URLs typically finish in 2 to 3 minutes, so an agent asked to compare 20 creators gets an answer in one tool call. Comments need a TikTok browser cookie passed in the cookie field; videos and profiles do not. - MCP endpoint: https://mcp.apify.com/?tools=fetch-actor-details,openclawai/tiktok-douyin-bilibili-scraper - REST: POST the input JSON to the run-sync-get-dataset-items endpoint and receive rows in the response body - Every row has item_type, so error handling is a field check, not an exception - Residential proxy is on by default; custom proxies go in proxyConfiguration.proxyUrls - 100 URLs in roughly 2 to 3 minutes per the actor documentation, August 2026 ### Steps 1. Open https://apify.com/openclawai/tiktok-douyin-bilibili-scraper, set mode (video_detail for one video, user_posts for a creator's feed, user_profile for account stats), and paste a TikTok URL or share link into url, or a list into urls. 2. Set maxItems (default 100, up to 5,000) for paginated modes, leave downloadVideos on for no-watermark MP4 links, and paste a TikTok browser cookie into the cookie field only if you are running video_comments. 3. Start the run and read rows from the dataset as JSON or CSV, or send the same input to the Apify REST API's run-sync-get-dataset-items endpoint to get the rows back in one HTTP response. 4. For agents, connect to https://mcp.apify.com/?tools=fetch-actor-details,openclawai/tiktok-douyin-bilibili-scraper and call the actor as a tool; mix Douyin and Bilibili URLs into the same urls array when you need cross-platform data. ### FAQ Q: Does TikTok have an official API for pulling public videos and comments? A: TikTok's Research API exists, but access is limited to academic institutions in the US, EEA, UK, or Switzerland, not-for-profit research bodies in the EU, and Brazilian institutions studying youth safety. The research must be non-commercial, pass an ethics review, and disclose its funding, and TikTok says to expect a reply within about 4 weeks (developers.tiktok.com, August 2026). Marketing and research teams without that access use a scraper like Datapika, which needs no developer account. Q: Why does TikTok comments mode return 0 results? A: TikTok serves comments only to logged-in sessions, so video_comments on TikTok needs a fresh browser cookie pasted into the cookie field, which is stored as a secret input. Videos, profiles, user posts, and likes work without one, and Bilibili comments run without a cookie. If the cookie is stale you get an error row, not a charge, so rotate it before large comment jobs. Q: Do the no-watermark TikTok MP4 links expire? A: Yes. TikTok's CDN links in video_url_nowm expire within a few hours, so download and store the file promptly rather than saving the URL. downloadVideos is on by default and works for TikTok and Bilibili. On Douyin the CDN currently rate-limits these links to 403 responses, so the option is auto-disabled there and metadata is still returned; forceDownload: true overrides that if your setup can fetch them. Q: Can I scrape TikTok by hashtag, keyword, or trending page? A: Not directly. The actor is URL-driven, and trending mode is limited to Douyin hot search and Bilibili popular, so there is no TikTok trending feed. The workaround is to run user_posts on a set of trend-setting creators and filter the hashtags array or music_title in your own code. If hashtag search is the core requirement, clockworks/tiktok-scraper accepts hashtag input at $1.70 per 1,000 results (verified August 2026). Q: Can I scrape private TikTok accounts or region-blocked videos? A: No. Only public content is returned; private, deleted, or region-blocked URLs come back as an error row with item_type set to error and a readable error_message, and those rows are not billed. Some regions soft-block TikTok profile lookups, in which case switching the residential proxy to another country group usually clears it. Custom proxies can be set in proxyConfiguration.proxyUrls. Q: Can I scrape TikTok, Douyin, and Bilibili in the same run? A: Yes, and that is the main difference from single-platform TikTok scrapers. Leave platform on auto, put tiktok.com, douyin.com, and bilibili.com URLs in the same urls array, and every row comes back with the same field names plus a platform tag. One caveat: the mode applies to the whole batch, so a user_posts run needs profile URLs for all three sites. Douyin user_posts can lag by about 6 days for some accounts. ## Turn job postings into hiring signals: one API for recruiters and sales Canonical: https://datapika.com/use-cases/hiring-signals-api Updated: 2026-08-29 A hiring signals API takes a list of company names and returns, for each one, whether it is hiring right now and how hard. Datapika's version checks Indeed and ZipRecruiter in a single run and returns a has_active_postings flag, employer-verified posting counts per board, the company's live ZipRecruiter job total, a firmographic profile, and its newest postings with hour-precise ZipRecruiter timestamps and repost detection. It costs $0.0005 per company record on Apify as of August 2026, so 1,000 companies come to $0.50 and 10,000 to $5, and the same call works over REST or MCP. ### What does a hiring signals API return? Each company you submit comes back as one record with the signal, the volume behind it, and the context you need to act. The signal is a boolean, has_active_postings, backed by active_jobs_by_site, which counts employer-verified postings on Indeed and ZipRecruiter from this run. The volume is active_jobs_total_by_site.zip_recruiter, the company's real open-posting total read from its ZipRecruiter employer page rather than a page-limited sample. The context is a profile block: industry, employee band, revenue band, headquarters, founding year, employee rating, website and socials. A sample record for a human resources company in the actor README shows 17 Indeed postings, 21 ZipRecruiter postings, a live ZipRecruiter total of 593 and a 7.27 employee rating. Nested under each record are the latest jobs, each with a posted_at timestamp (hour-precise on ZipRecruiter, day-precise on Indeed) and a separate reposted_at field where the employer bumped an old ad. Runs default to a 48-hour window and 50 jobs per company per board. - has_active_postings: the yes/no hiring signal per company - active_jobs_by_site: employer-verified posting counts for Indeed and ZipRecruiter - active_jobs_total_by_site.zip_recruiter: the full live ZipRecruiter total, a hiring-volume number - posted_last_48h_count: postings inside your recency window, default 48 hours - latest_jobs: title, salary range, location, remote flag, posted_at, reposted_at and application URL ### How do recruiters use hiring signals to find clients who are hiring? Staffing and agency recruiters run their client and prospect list through companies mode each morning after the boards refresh. Any account whose posted_last_48h_count moves above zero enters the outreach queue with the job titles, salary ranges and locations attached, so the first line of the email names the role the hiring manager opened yesterday. Repost detection matters more here than anywhere else. When an employer keeps bumping the same ad, the record keeps the original posting date and puts the bump in reposted_at, a reliable sign of a hard-to-fill role and a stronger staffing pitch than a new opening. The profile block handles qualification in the same pass. At $0.0005 per record, a daily scan of 2,000 client companies costs $1 and returns the fresh roles alongside the signal, so there is no second enrichment step. - Schedule companies mode daily and act on rows where posted_last_48h_count went positive - Add the website after the name (Kelly Services | kellyservices.com) so common names match the right employer - Use reposted_at to spot relisted roles, often the hardest to fill - Qualify with industry, employee band, revenue band, HQ and rating from the same record - Feed the flat job rows into a candidate-matching pipeline while the posting is hours old ### How do sales teams turn hiring intent data into account scores? Hiring is spend that has already been approved, which makes it one of the earliest buying signals visible from outside a company. A team opening five engineering roles is about to buy laptops, cloud capacity, licenses and services for them, and a rep who knows that this week is ahead of any announcement. The company scan turns a target account list into per-account intent built from live numbers. The practical pattern is to map two fields into the CRM: the live ZipRecruiter total as a size-of-expansion field, and posted_last_48h_count as a velocity field. Sort by velocity for this week's outreach and by total for territory planning. Industry and revenue segmentation come from the same record. Rescanning is cheap enough to do weekly. A 5,000-account list costs $2.50 per pass, so 13 weekly scans over a quarter come to $32.50. - Score accounts by live ZipRecruiter total (size) and posted_last_48h_count (velocity) - Route accounts whose count jumped week over week into a fresh sequence, not the generic cadence - Use role titles as the hook: a first sales hire in a new region, or a cluster of platform roles - Export CSV or call the dataset API and write both counts into custom intent fields ### How do analysts build a companies-hiring dataset for competitive intel? Analysts want a time series, not a snapshot. Schedule the same company list on Apify every week and each run writes a dated dataset with posting counts, live totals and the newest titles per competitor, so the week-over-week diff becomes the report. The guide to tracking competitor hiring covers the scheduling and diffing pattern in detail. Keywords mode covers the market view. Up to 10 search terms per run return fresh postings across both boards with salary_min, salary_max, interval and currency, plus a per-term summary of totals by board and how many fell inside the window. That gives posting volume and pay ranges per role and region without a company list. Profiles mode is the lightweight enrichment path: names in, industry, size, revenue, HQ, founding year and rating out, with no job feed attached. - Companies mode on a weekly schedule produces a dated hiring time series per competitor - Keywords mode: up to 10 terms per run, salary normalized to min, max, interval and currency - Profiles mode: firmographics only, for fast enrichment of large lists - Indeed supports country selection; ZipRecruiter covers the US and Canada - Export JSON, CSV or Excel, or push each run to Make, Zapier or Google Sheets via webhooks ### Steps 1. Open https://apify.com/openclawai/indeed-ziprecruiter-scraper, leave scanMode on companies, and paste your company list, adding the website after a pipe for exact matching. 2. Set hoursOld (default 48) and maxJobsPerCompany (default 50), keep residential proxy on, and run. One record per company lands in the dataset with the signal, counts, profile and latest jobs. 3. Export the dataset as CSV or JSON, or pull it through the Apify API, and route rows with has_active_postings true into your outreach or scoring pipeline. 4. For agents, connect the MCP endpoint at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/indeed-ziprecruiter-scraper and call the actor with the same input from Claude, Cursor or your own agent. ### FAQ Q: Does Indeed or ZipRecruiter have an official hiring signals API? A: Not one you can sign up for and start calling today. Both boards run developer programs built for employers, applicant tracking systems and job publishers rather than for reading which companies are hiring. Datapika reads the public job and employer pages instead, merges both boards into one dataset, and exposes the result through the Apify REST API and an MCP endpoint, billed per company record. Q: How much does hiring intent data cost with Datapika? A: Company records cost $0.0005 each on Apify as of August 2026. That is $0.50 for 1,000 companies, $5 for 10,000, and $2.50 for a weekly pass over a 5,000-account book. You pay per result with no subscription, and the price covers the hiring signal, per-board counts, live ZipRecruiter total, company profile and latest postings for every company on the list. Q: How fresh is the hiring signal, and how does repost detection work? A: ZipRecruiter rows carry hour-precise UTC posted_at timestamps read from the page itself, while Indeed rows are day-precise. When an employer bumps an old ad, the record keeps the original date and reports the bump separately in reposted_at, so a 24-hour list contains only genuinely new roles. One caveat: ZipRecruiter ingests postings in daily batches, so its newest jobs are usually 12 to 48 hours old; use a 48-hour window for full two-board coverage. Q: How accurate is company matching for common names? A: Jobs are matched to the target company by website domain first, then by normalized name as a fallback. If you supply the website after a pipe, for example Kelly Services | kellyservices.com, generic names such as Kelly or Volt will not pull in unrelated employers. The active_jobs_by_site counts are employer-verified for this reason, and the same match drives the live ZipRecruiter total. Q: Can an AI agent call the hiring signals API directly? A: Yes. The actor is exposed through the Apify MCP server with the fetch-actor-details tool, so an agent can read the input schema, run a company scan and read the dataset without a human in the loop. Billing stays per result at $0.0005 per company record, which suits agent workflows that check a handful of accounts on demand rather than bulk exports. ## A scraping API your agent can call, and pay for, per result Canonical: https://datapika.com/use-cases/scraping-api-for-ai-agents Updated: 2026-08-29 Datapika is a pay-per-result scraping API for AI agents: 10 actors on the Apify platform, each callable as an MCP tool or a REST endpoint, billed per delivered row with no subscription. As of August 2026 prices start at $0.0005 per record on the hiring signals scanner and reach $0.005 per job on the flagship job board actor. One MCP URL preloads every tool into Claude Code, Cursor, VS Code or Codex CLI, and a maxTotalChargeUsd cap fixes a hard budget before the call. Plain-text llms.txt, AGENTS.md and per-actor OpenAPI specs let the agent read inputs and pricing itself. ### What does a pay-per-result scraping API cost an AI agent? Every Datapika run bills per delivered result and nothing else. There is no monthly plan, no seat and no minimum: a free Apify account is enough to start, and that account is charged for the rows an actor returns. The rate is set per actor. As of August 2026 the Apify Store lists the hiring signals scanner at $0.0005 per record, the Google Flights actor at $0.0015 per itinerary, the Reddit actor at $0.002 per post, the Google Maps actor at $0.003 per place and the multi-board job actor at $0.005 per job. Trend reports are priced per report instead, at $0.35 for the standard tier. Because the unit is a row, an agent can compute the cost of a call before making it and skip calls that exceed its allowance. - A 100-job sweep on the job board actor costs $0.50; 1,000 hiring-signal company checks also cost $0.50 - 1,000 Google Flights itineraries cost $1.50; a 1,000-place Google Maps run is $3.00 at a flat $0.003 per place - maxTotalChargeUsd on any REST run stops the platform at that ceiling instead of overspending - The TikTok actor returns an unbilled error row for a failed URL, and the Airbnb actor charges nothing for empty or failed runs, so retries do not compound cost ### How does an AI agent connect to Datapika over MCP? The actors are exposed through the Apify MCP server over streamable HTTP with OAuth. One URL preloads every Datapika tool by listing each slug after fetch-actor-details, for example https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper,openclawai/reddit-scraper and so on through the catalog. Each actor also has a minimal single-tool URL of the form https://mcp.apify.com/?tools=fetch-actor-details,openclawai/, which is the right choice when an agent needs only one source. The paired fetch-actor-details tool is what makes the setup agent-safe. Before running anything, the model can read the actor's input schema and current price and decide whether the call fits the budget it was given. On first use the client opens Apify OAuth in the browser; a free account completes it, and every run after that bills per result to the same account. - Claude Code: claude mcp add --transport http datapika "" - Cursor: the Add to Cursor button on datapika.com/mcp, or the URL in .cursor/mcp.json - VS Code: .vscode/mcp.json with type http; Codex CLI: codex mcp add datapika --url "" - Example prompt: find remote senior data engineer roles posted in the last 24 hours, dedupe across boards, return salary ranges as a table ### When should an agent use the REST API instead of MCP? Use REST when the agent runs unattended, on a schedule, or inside a pipeline that has no MCP client. A single POST to https://api.apify.com/v2/acts/openclawai~/run-sync-get-dataset-items?format=json with a Bearer token returns dataset rows directly as JSON. The synchronous endpoint returns 408 after 300 seconds, so large sweeps should start a run, poll its status, then read the dataset when it finishes. Official Node and Python clients wrap both patterns, and the Apify CLI covers shell scripts. Each actor publishes an OpenAPI spec for its default build at https://apify.com/openclawai//api/openapi with no auth required, so codegen or a tool-calling framework can build a typed client from it. Schedules and webhooks are platform features: run any actor every morning and push new rows to your stack. - Sync: run-sync-get-dataset-items returns rows in one response for runs that finish inside 300 seconds - Async: actor.call() then dataset.listItems() in the Node or Python client, required past the 300-second limit - Budget: append maxTotalChargeUsd= to the run URL to cap spend per call - Auth: an Apify API token in an Authorization: Bearer header, issued from a free account ### How does an agent discover Datapika's tools without a human? Three plain-text surfaces describe the catalog in a form a model can read directly. datapika.com/llms.txt follows the llms.txt convention: a one-paragraph description of the studio, every actor with its price and run URL, the MCP config URL, the guides, and a facts block. datapika.com/llms-full.txt carries every guide and actor description in one file for retrieval. datapika.com/AGENTS.md uses the installation, configuration and usage layout that coding agents already expect from repository AGENTS.md files: the MCP add command, the Bearer auth rule, the budget parameter, the 300-second timeout, and one line per actor with slug, price, docs link, input schema link and OpenAPI link. All three render from the same catalog data as the pages, so a price never drifts between what a human reads and what an agent reads. - llms.txt: catalog map with per-actor price and run URL, plus the one-line MCP install - AGENTS.md: installation, configuration and usage sections, including timeouts and budget caps - Input schema at https://apify.com/openclawai//input-schema; OpenAPI at https://apify.com/openclawai//api/openapi - Pages, sitemap, llms.txt and AGENTS.md share one data source, so a price change updates everywhere at once ### Steps 1. For an interactive agent, add the MCP server: claude mcp add --transport http datapika "https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper" (or the full catalog URL from datapika.com/mcp for all 10 tools), then complete Apify OAuth once with a free account. 2. For unattended pipelines, POST the actor input to https://api.apify.com/v2/acts/openclawai~job-board-scraper/run-sync-get-dataset-items?format=json&maxTotalChargeUsd=1 with an Authorization: Bearer token; the response body is the dataset as JSON. 3. Let the agent read datapika.com/AGENTS.md or call fetch-actor-details to learn each actor's input fields and per-row price, and set maxResults to the rows it actually needs. 4. Read the rows back as JSON. If a sweep may pass 300 seconds, start it asynchronously and poll, then fetch the dataset; add a schedule and webhook for recurring pulls. ### FAQ Q: Does Datapika require a subscription or a minimum spend? A: No. Every actor bills per delivered result to an Apify account, which is free to create and needs no card to start. You pay the per-row price of each actor you call, from $0.0005 per record on the hiring signals scanner to $0.005 per job on the multi-board job actor, as of August 2026. Trend reports are priced per report instead, at $0.35 for the standard tier. Q: Can an AI agent pay for a scraping call without a human-created account? A: Not reliably yet. Runs bill against an Apify account today, so a person creates the account and token once and the agent uses them. The platform is rolling out agentic payment protocols, x402 and MPP prepaid tokens, that would let an agent buy usage directly; support varies by actor and account status, so check the platform docs before relying on it. Until then, an agent-held token plus maxTotalChargeUsd is the working pattern. Q: What happens when a scraping run exceeds the 300-second synchronous limit? A: The run-sync-get-dataset-items endpoint returns HTTP 408 after 300 seconds. The run is not the problem, the open connection is: switch to the asynchronous pattern, start the run, poll its status, then read the dataset items once it finishes. The official Node and Python clients do this with actor.call() followed by a dataset listItems() read. Large multi-board job sweeps and full comment-tree Reddit pulls are the usual cases. Q: How does an agent stop a scraping run from overspending? A: Two controls. maxTotalChargeUsd is a query parameter on any REST run, and the platform stops the run when charges reach that ceiling. maxResults, the per-actor input, limits how many rows the actor tries to return in the first place. Because the MCP setup pairs every tool with fetch-actor-details, an agent can also read the price per row before it runs and multiply: a 100-job sweep at $0.005 per job is $0.50. Q: Why not call the official APIs of these sites from an agent instead? A: Most of the sources Datapika covers offer no public API for this data to independent developers, and where an official API does exist it needs its own key, quota and billing arrangement. Datapika reads public pages and returns normalized JSON, so the agent uses one auth and one price model across every source: one Apify token, per-row pricing, and the same sync or async run pattern whether the target is a job board, Reddit, Google Maps or Google Flights. ## Track trends across Reddit, TikTok, and YouTube from one topic query Canonical: https://datapika.com/use-cases/trend-monitoring Updated: 2026-08-29 Datapika's trend intel actor runs a single topic query across 13+ platforms, including Reddit, TikTok, Instagram, YouTube, Hacker News, GitHub, Polymarket, Threads, Pinterest, and Bluesky, over a window of 7, 14, 30, or 90 days. Every result is scored by real engagement (upvotes, views, likes) and merged into cross-source clusters, so the same story on Reddit and YouTube becomes one insight. From August 31, 2026 the actor is free to run, with only your Apify platform compute usage billed; you still choose a report depth from Quick Scan to Full Brief, with Standard as the default. Optional AI synthesis adds an executive summary, key findings, sentiment, and momentum, delivered as JSON or Markdown with no API keys required. ### How does cross-platform trend analysis work from a single query? You enter one topic, such as a product name, a person, a company, or a phrase like "best CRM tools 2026", pick a timeframe, and optionally restrict the sources. The actor searches every platform in your tier and returns one dataset instead of ten separate exports. Ranking uses the engagement the platforms themselves report, so a Reddit thread with 2.1K upvotes outranks a thin page that merely ranks well in search. Cross-source clustering is what saves analyst time. When the same story surfaces on Reddit, Hacker News, and YouTube, the actor merges those hits into one cluster and reports how many platforms carried it. The README's example output for "OpenAI vs Anthropic" shows 47 results across six sources: 15 from Reddit, 10 from Hacker News, 8 from TikTok, 6 from YouTube, 5 from GitHub, and 3 from Polymarket. - Input: one topic string, a timeframe of 7d, 14d, 30d, or 90d, and an optional list of sources - Per-platform payloads: Reddit posts with upvotes and subreddit breakdown, TikTok and YouTube videos with view counts and transcripts, GitHub repos with stars and releases - Dataset row fields: total_results, source_counts per platform, cluster_count, tier, price, and generated_at, with ranked_candidates and clusters in the full report - Polymarket adds prediction-market odds backed by real money, useful for events and launches - Every run fetches live data, so a report reflects today's engagement numbers ### Which report tier should you use for trend monitoring? The four tiers differ in which platforms are queried and how much is pulled per platform, not in output format. Quick Scan covers Reddit, Hacker News, Polymarket, and GitHub, which is enough for a daily pulse on technical or financial topics. Standard adds TikTok, Instagram, and YouTube and is the tier most teams should start with. Deep Intel keeps the same platforms but fetches more results per source plus comment threads and video transcripts, which matters when you want quotes rather than headlines. Full Brief adds AI web grounding so the synthesis can cite pages outside the ten social sources. As of August 2026 the actor has logged 215 runs on the Apify Store. From August 31, 2026 there is no per-report charge for any tier: a run costs only the Apify platform compute it consumes, and deeper tiers use more compute because they fetch more. - Quick Scan: free-tier sources only, best for high-frequency checks - Standard: adds the three short-video and image platforms where consumer trends start - Deep Intel: more results per source plus comments and transcripts for quote mining - Full Brief: everything above plus web grounding for the AI summary - No per-report fee from August 31, 2026; the only cost is Apify compute usage, so tier choice is about depth, not price ### What does the AI synthesis add to a trend report? With aiSynthesis set to true (the default), the actor writes an executive summary over the clustered results and extracts key findings, each tagged with its source platform and engagement figure, for example a finding sourced from Reddit with 2.1K upvotes. The README's output example also classifies overall sentiment as bullish, bearish, neutral, or mixed, and momentum as accelerating, steady, declining, or emerging. The momentum label is the field that makes the actor useful for monitoring rather than one-off research. Run the same topic weekly and the momentum value tells you whether attention is building before the raw counts make it obvious. If the synthesis step fails for any reason, the run still returns the full scored and clustered dataset and, from August 31, 2026, costs nothing beyond the compute the run used. - executive_summary: a short narrative over the top clusters, written for a reader who has not seen the data - key_findings: an array of findings, each with a source platform and an engagement number - sentiment: one of bullish, bearish, neutral, or mixed - momentum: one of accelerating, steady, declining, or emerging - best_takes: the highest-engagement quotes across platforms, useful for content and sales prep ### How do you track trends across Reddit, TikTok, and YouTube on a schedule? Schedule the actor on Apify with a fixed topic and a 7d timeframe, and each run becomes one row in a time series. From August 31, 2026 the actor itself is free, so a weekly Standard report or a daily Quick Scan costs only the Apify platform compute each run consumes, which makes high-frequency monitoring on several topics practical. Choose JSON output when the rows feed a dashboard or a database, and Markdown when the report is emailed to people. For agents, the same actor is exposed through Apify's MCP server, so an assistant can call it by name, read the momentum and sentiment fields, and decide whether to alert a human. The sources filter lets you narrow a scheduled run to the platforms that matter for your audience, for example TikTok and Instagram for a consumer brand, or Hacker News and GitHub for a developer tool. - Four weekly Standard runs per topic per month, with momentum visible after the second run and no per-report fee from August 31, 2026 - Daily Quick Scan: 30 runs per month per topic, billed only as Apify compute usage, and the lightest tier keeps that compute low - Sources filter accepts any subset of reddit, tiktok, instagram, youtube, hackernews, polymarket, github, threads, pinterest, and bluesky - Residential proxy is on by default so scheduled runs are not blocked by platform 403 responses - MCP endpoint lets an agent run a report and read the result in one tool call ### Steps 1. Open https://apify.com/openclawai/30days-trend-intel, enter a topic, and pick a timeframe of 7, 14, 30, or 90 days. 2. Choose a tier (Quick Scan, Standard, Deep Intel, or Full Brief), leave AI synthesis on, and optionally limit the sources list; from August 31, 2026 no tier carries a per-report fee. 3. Run it once to check the clusters and momentum label, then save the input as a schedule on Apify to build a weekly or daily series. 4. For agents, connect https://mcp.apify.com/?tools=fetch-actor-details,openclawai/30days-trend-intel and call the actor by name with the same input. ### FAQ Q: How much does cross-platform trend monitoring cost with Datapika? A: From August 31, 2026 the actor is free to run: there is no per-report charge for any tier, and you pay only for your own Apify platform usage, meaning the compute a run consumes plus any residential proxy bandwidth, which Apify bills at its standard rates. No platform API keys or subscriptions are needed. Deeper tiers such as Deep Intel and Full Brief fetch more per source and therefore use more compute than a Quick Scan, so pick the lightest tier that answers your question for high-frequency schedules. Q: Do Reddit, TikTok, and YouTube have official APIs for trend tracking? A: Each platform has its own developer API with separate keys, quotas, and approval steps, and a self-built tracker breaks whenever one of them changes its terms or rate limits. Datapika's actor removes that setup: one input covers all sources, credentials are handled by the actor, and residential proxies are on by default. You still get engagement figures the platforms themselves report, such as upvotes, views, and likes. Q: Which platforms are included, and can I limit the report to a few of them? A: The README lists 13+ platforms searched, and the sources filter exposes ten by name: Reddit, TikTok, Instagram, YouTube, Hacker News, Polymarket, GitHub, Threads, Pinterest, and Bluesky. Leave the filter empty to search everything your tier allows, or pass a subset such as tiktok and instagram for a consumer brand. Note that TikTok, Instagram, and YouTube require the Standard tier or above. Q: How fresh is the data in a trend report? A: Every run fetches live data at execution time, so there is no cached index behind the results. The timeframe you pick (7, 14, 30, or 90 days) controls how far back each platform search reaches, and generated_at records the exact run time. For monitoring, a 7-day window on a weekly schedule gives non-overlapping slices, while a 30-day window smooths out single viral spikes. Q: What happens if the AI synthesis step fails? A: You still receive the complete raw dataset: ranked candidates with engagement scores, cross-source clusters, and per-platform counts. The executive_summary, key_findings, and best_takes fields come back empty, or reduced to a plain extraction of the top-ranked titles, and from August 31, 2026 the run costs only the Apify compute it used. A pipeline should therefore treat the synthesis fields as optional and rely on the scored results as the source of truth. ## Indeed shut its public API. Here's how to get job data anyway Canonical: https://datapika.com/compare/indeed-api-alternative Updated: 2026-08-29 Indeed has no public job search API in 2026: the Publisher Program closed to new publishers in October 2022, and the partner APIs that remain only push postings into Indeed, never read them out. For the cheapest Indeed-only feed, valig/indeed-jobs-scraper lists $0.10 per 1,000 jobs with a 5.00 rating and 26,654 users. For the longest track record, misceres/indeed-scraper costs $3.00 per 1,000 at 3.76 stars and 29,854 users. For Indeed plus 7 more boards deduplicated in one run, Datapika's job board scraper costs $0.005 per job ($5 per 1,000). All figures as of August 29, 2026. ### Does Indeed have an API? What happened to the Publisher API? No. Indeed has not offered a public job search API since it closed the Publisher Program to new applicants in October 2022, and the program remains closed in 2026 with no open signup and no published rate card (Job Boardly, July 22, 2026). Existing Publisher API keys stopped working in 2023, and Indeed has issued no self-serve key since, paid or free (JobsPipe, July 24, 2026). What remains at docs.indeed.com is a set of employer-side partner APIs, all gated behind an integration request that Indeed must approve before it issues OAuth credentials. None of those partner APIs returns job listings. The Job Sync API is a GraphQL interface for ATS partners to create, upsert, expire, and check the status of their own postings by ID. The only search-shaped product Indeed still documents for publishers is the Publisher JavaScript Plugin, a hosted widget that renders HTML inside your page rather than returning data. If your goal is to read jobs from Indeed, the official route does not exist, which is why every option below is a scraper. - Job Sync API: partners push postings into Indeed with qualifications, salary, and benefits; status lookups need the posting ID, and there is no search endpoint - Indeed Apply: delivers applications from Indeed into a partner ATS; the job must be posted through Job Sync or an XML feed, and Indeed must approve an integration request first - Disposition Sync API: reports candidate outcomes back to Indeed, no listing retrieval - Publisher JavaScript Plugin: a hosted search widget that returns rendered HTML, not JSON - Publisher Program: closed to new publishers since October 2022; existing keys retired in 2023 (sources dated July 2026) ### Indeed API alternatives compared: price per 1,000, rating, users Five scrapers cover the read-Indeed use case, and the table records their live Apify Store numbers as of August 29, 2026. The spread is wide: valig/indeed-jobs-scraper headlines $0.10 per 1,000 jobs (its pricing page shows from $0.07 per 1,000) with a 5.00 rating and 26,654 users, while borderline/indeed-scraper charges $5.00 per 1,000 at 4.63 stars and 22,604 users. misceres/indeed-scraper is the incumbent with 29,854 users and 2,299 in the last month, yet it carries the lowest rating on the table at 3.76 and costs $3.00 per 1,000. factden/indeed-jobs-scraper sits at $2 per 1,000 (down to $1.20 on higher Apify plans) with a 5.00 rating but only 44 users, so its score rests on very little public evidence. All five bill per event, so you pay for rows delivered rather than compute time. Datapika's job board scraper is the only multi-source row: $0.005 per job ($5 per 1,000), a 5.0 rating from 3 reviews, 2,471 users with 381 in the last 30 days, and a 100% run success rate. It is not the cheapest Indeed-only feed. It wins on covering Indeed plus LinkedIn, Glassdoor, Google Jobs, ZipRecruiter, Naukri, Bayt, and BDJobs in one deduplicated dataset, with no API key or login. - Cheapest Indeed-only feed: valig at $0.10 per 1,000 (5.00 rating, 26,654 users, 100% success) - Most used: misceres at $3.00 per 1,000 (3.76 rating, 29,854 users, 99.7% success) - Highest-priced single board: borderline at $5.00 per 1,000 (4.63 rating, 22,604 users, 97.8% success) - Newest: factden at $2 per 1,000 (5.00 rating, 44 users, 100% success) - Multi-board with dedup and MCP: Datapika at $5 per 1,000 (5.0 rating, 2,471 users, 100% success) ### What does it cost to pull 1,000 or 10,000 Indeed jobs? Pay-per-event pricing keeps the arithmetic simple: multiply rows by the per-row price, and the Apify platform charge is zero for all five actors. For 1,000 Indeed jobs, valig costs $0.10 (1,000 x $0.0001), misceres costs $3.00 (1,000 x $0.003), and Datapika costs $5.00 (1,000 x $0.005). At 10,000 jobs that becomes $1.00, $30.00, and $50.00. The Datapika number changes meaning when you need more than Indeed. A 10,000-row sweep across Indeed, LinkedIn, and Glassdoor with three single-board actors is three separate runs, three schemas to reconcile, and duplicate postings you pay for more than once. The same sweep through Datapika is one run at $0.005 per deduplicated row, so a job that appears on all three boards is billed once and tagged with every source. Compare like with like. If Indeed is your only source, valig at $1.00 per 10,000 is 50x cheaper than Datapika and the honest recommendation. If you already run two or more boards, price the whole pipeline, not the Indeed slice. - 1,000 jobs: valig $0.10, factden $2.00, misceres $3.00, borderline $5.00, Datapika $5.00 - 10,000 jobs: valig $1.00, factden $20.00, misceres $30.00, borderline $50.00, Datapika $50.00 - Datapika bills only deduplicated rows that land in the dataset; a posting seen on three boards costs $0.005 once - No platform compute charge on any of the five; the per-row price is the whole bill (Apify Store pricing pages, August 29, 2026) ### Field coverage: salary, company data, descriptions, deduplication Every actor on the table returns title, company, location, URL, and a description. The differences are in the structured extras. valig parses salary min and max, benefits such as paid holidays and 401(k), required skills, and job type, and returns descriptions as text and HTML. borderline adds company rating, company logo, emails, latitude and longitude, and urgent-hire and high-volume flags. factden ships a second companies dataset with CEO, founding year, revenue, and social links, and documents date-window sharding past Indeed's roughly 1,000-result depth limit per query. Datapika returns salary min, max, currency, and interval with optional annual normalization, company size, revenue, rating, review count, industry, and logo, and the full description in markdown or HTML when fetch-description is on. Its distinctive field is the source tag on every row, which only matters because a run can span 8 boards. For Indeed-only extraction, field depth is roughly a wash across the top four; for cross-board work, the source tag and deduplication are the reason to pay more per row. - Structured salary min and max: valig, borderline, factden, and Datapika all parse it; Datapika adds currency, interval, and annual normalization - Company enrichment: borderline (rating, emails), factden (separate companies dataset), Datapika (size, revenue, rating, reviews, industry, logo) - Deep pagination past Indeed's ~1,000-result cap: factden documents it; the others do not advertise it - Cross-board source tags and deduplication: Datapika only, across 8 boards - Agent access: Datapika ships a ready MCP link (mcp.apify.com) and needs no API key or login; single-board actors can also run through the Apify API or MCP ### Which Indeed API alternative should you choose? Choose valig when Indeed is your only source and price dominates: at $0.10 per 1,000 with a 5.00 rating and 3,627 monthly users, it is the default for budget feeds. Choose misceres if you want the longest track record; 29,854 users and 99.7% success are real, but budget for the 3.76 rating and $3.00 per 1,000. Choose factden if you must paginate past 1,000 results per query or want the companies dataset, accepting that 44 users means limited public evidence. Choose Datapika when the question is not how to read Indeed but how to get every relevant job once. Recruiters, hiring-signal pipelines, and agents that search across boards get one schema, one bill at $0.005 per row, and no keys. Skip it if you will only ever query Indeed; at 50x the price of valig for that single board, it is the wrong tool. If you need to post jobs into Indeed rather than read them, none of these applies: file an integration request with Indeed and use the Job Sync API. - Indeed-only, lowest cost: valig ($0.10 per 1,000) - Indeed-only, longest track record: misceres ($3.00 per 1,000, 29,854 users) - Indeed-only, deep pagination and company profiles: factden ($2 per 1,000) - Indeed plus 7 boards, deduplicated, MCP-ready: Datapika ($5 per 1,000, 5.0 rating) - Posting jobs into Indeed rather than reading them: request partner access and use the Job Sync API ### Steps 1. Open https://apify.com/openclawai/job-board-scraper, enter up to 5 search terms, tick Indeed plus any of the other 7 boards, and set the country for Indeed. 2. Add filters you need (remote only, job type, posted within N hours, easy apply) and turn on fetch-description if you want full posting text in markdown or HTML. 3. Start the run. First rows arrive in about 3 seconds, you are billed $0.005 per deduplicated job delivered, and you can export JSON or CSV or read the dataset through the Apify API with no Indeed credentials. 4. For agents, register https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper in Claude, Cursor, or any MCP client and call the scraper as a tool with the same inputs. ### FAQ Q: Does Indeed have a public API in 2026? A: No. Indeed's partner documentation covers Job Sync (post and manage your own listings via GraphQL), Indeed Apply (receive applications), and Disposition Sync (report candidate outcomes). All require an approved integration request and OAuth credentials issued by Indeed, and none returns job postings you did not create. The Publisher JavaScript Plugin embeds a hosted search widget but returns rendered HTML, not data. If you want to read listings rather than publish them, there is no official endpoint, as of August 2026. Q: Can I still get an Indeed Publisher API key? A: No. The Publisher Program closed to new publishers in October 2022 and remains closed in 2026 (Job Boardly, July 22, 2026). Existing keys stopped working in 2023 when Indeed retired the API in favour of a hosted search widget, and no self-serve key has been issued since, paid, free, or waitlisted (JobsPipe, July 24, 2026). Guides that still show a publisher-key example are describing an endpoint that no longer answers. Q: Is it legal or allowed to scrape Indeed job listings? A: Scraping publicly visible listings for analysis is common, and every actor in this comparison runs on Apify under its acceptable-use terms. Indeed's own terms prohibit automated access, so scraping is not an Indeed-sanctioned route, and Indeed's bot protection can block IPs. Personal data inside postings, such as recruiter names or emails, falls under GDPR and similar laws and needs a lawful basis. This is not legal advice; check your jurisdiction before building a product on scraped data. Q: Which Indeed scraper is cheapest per 1,000 jobs? A: valig/indeed-jobs-scraper, at a headline $0.10 per 1,000 jobs with its pricing page showing from $0.07 per 1,000, as of August 29, 2026. It also carries a 5.00 rating, 26,654 users, and a 100% run success rate on the store. factden is next at $2 per 1,000, misceres at $3.00, and borderline and Datapika at $5.00. All five are pay-per-event, so the per-row price is the entire cost with no separate compute charge. Q: Why pay $5 per 1,000 for Datapika when valig costs $0.10? A: You should not, if Indeed is your only source. Datapika's $0.005 per job buys a single run that sweeps Indeed and 7 other boards (LinkedIn, Glassdoor, Google Jobs, ZipRecruiter, Naukri, Bayt, BDJobs), deduplicates postings across them, tags each row with every source, and returns salary and company fields in one schema. Priced against three single-board actors plus your own dedup code, the comparison changes. Priced against valig alone, valig wins by 50x. Q: Can an AI agent use an Indeed API alternative without an API key? A: Yes. Datapika's job board scraper publishes an MCP endpoint at mcp.apify.com, so Claude, Cursor, or any MCP client can call it as a tool: pass search terms, boards, and country, and receive deduplicated rows at $0.005 each. No Indeed login, cookies, or publisher key is involved, and billing is per delivered row. Single-board actors on Apify can also be reached through the Apify API or MCP; the difference is that one Datapika call covers 8 boards. ## Which Google Flights scraper should you use? Ten options compared Canonical: https://datapika.com/compare/best-google-flights-scrapers Updated: 2026-08-29 Quick answer: there is no official Google Flights API, so every option reads the public results page. For proven volume at a low rate, memo23 at $0.001 per result plus $0.005 per run start (476 users, 99.9% success) is the value pick, and johnvc is the most used at 2,480 users with a 4.0 rating. Datapika's Google Flights scraper at $0.0015 per itinerary ($1.50 per 1,000, 3,696 runs, 93.5% success) fits when you need CO2 in grams, 30 or more result languages, multi-city trips and a hard per-query row cap. Skip the $19.99 a month rental at 0% success. ### What happened to the official Google Flights API? There is no Google Flights API you can sign up for. Google's only public airfare feed was QPX Express, built on the ITA Software technology it bought for $700 million in a deal that closed in 2011. On November 1, 2017 Google told TechCrunch it would retire the API "given the low interest among our travel partners", and it switched the service off on April 10, 2018. Before the shutdown the API cost $0.02 per query after 50 free queries a month, down from an earlier $0.035. The enterprise QPX product continued for large contracted travel companies, and a Google spokesperson said it generated the vast majority of the travel data revenue. Eight years later nothing public has replaced it, so every option on this page reads the Google Flights results page and converts what a traveler sees into rows. The tools differ in what they bill (per itinerary, per search, per results page, per month), how much of the page they return, and how often a run actually finishes. - QPX Express retirement announced November 1, 2017; shut down April 10, 2018 (TechCrunch) - Google's stated reason: low interest among travel partners - Last public price was $0.02 per query after 50 free queries a month - No public replacement as of August 29, 2026; enterprise QPX is contract-only - Every scraper below parses the public results page, so success rate is what separates them ### How do the best Google Flights scrapers compare on price, rating and users? Ten actors were checked on the Apify Store on August 29, 2026, and only three carry reviews: johnvc's Google Flights API at 4.0 stars from 6 reviews, memo23 at 5.0 from a single review, and makework36 at 1.0 from one review. The rest are unrated, so the table leans on the two numbers Apify publishes for every actor, total users and 30-day success rate. Johnvc has by far the largest base at 2,480 users and 377 in the last 30 days, but bills $0.03 per results page plus $0.01 per run on the free plan, the most expensive per search here. Memo23 is the value pick with volume behind it: $0.001 per result plus $0.005 per start, 476 users and 99.9% success. Automation-lab is cheaper still at $0.00023 per result on the free plan plus $0.005 per start, with 396 users and 100%. Kaix lists $0.06 per 1,000 on the free plan but bills Apify compute on top. Datapika charges $0.0015 per itinerary plus platform usage, with 3,696 lifetime runs and 93.5% success. Scrape.badger bills $0.012 per search, skootle $8 per 1,000 on the free plan, and ScrapeBase rents at $19.99 a month with 0% of runs succeeding. - Only rated actors: johnvc 4.0 (6 reviews), memo23 5.0 (1 review), makework36 1.0 (1 review) - Largest user base: johnvc, 2,480 users and 377 in the last 30 days - Cheapest listed rate with no compute bill: automation-lab at $0.23 per 1,000 on the free plan plus $0.005 per start - Highest success among actors with real volume: automation-lab and johnvc 100%, memo23 99.9%, Datapika 93.5% - Only per-search billing: scrape.badger at $0.012 per search, documented as an MCP tool - Only monthly rental: ScrapeBase at $19.99 plus usage, 0 monthly users and 0% success ### What does it cost to scrape 1,000 and 10,000 Google Flights itineraries? Figures use each actor's free-plan rate on August 29, 2026 and assume 25 itineraries per search, which is Datapika's default max_results and inside the 30 to 80 Google usually shows, so 1,000 itineraries is 40 searches. Datapika: 1,000 times $0.0015 is $1.50 and 10,000 is $15.00, plus Apify platform usage for the run time, which the listing does not fix in advance. Memo23: $1.00 in results plus 40 starts at $0.005 is $1.20; 10,000 is $10.00 plus $2.00, $12.00. Automation-lab: $0.23 plus $0.20 in starts is $0.43; 10,000 is $2.30 plus $2.00, $4.30. Scrape.badger: 40 searches at $0.012 is $0.48 and 10,000 is $4.80, but if you keep only the five cheapest fares per route the same 1,000 rows need 200 searches and cost $2.40. Johnvc at one results page per search: 40 pages at $0.03 plus 40 runs at $0.01 is $1.60, and 10,000 is $16.00. Skootle: $8.00 plus $0.20 is $8.20, and $82.00 for 10,000. Apify's own $0.00005 start and $0.00001 dataset-item accounting events add under 2 cents per 1,000 rows on any of them. - 1,000 itineraries: automation-lab $0.43, scrape.badger $0.48, memo23 $1.20, Datapika $1.50 plus compute, johnvc $1.60, skootle $8.20 - 10,000 itineraries: automation-lab $4.30, scrape.badger $4.80, memo23 $12.00, Datapika $15.00 plus compute, johnvc $16.00, skootle $82.00 - Per-search billing flips against you when you cap results low: 5 rows per search makes scrape.badger $2.40 per 1,000 - Kaix is left out of the arithmetic because its compute charge is not published on the listing - Paid Apify plans cut most rates: automation-lab drops to $0.12 per 1,000 and skootle to $5.00 on Gold ### Which fields and features differ between Google Flights scrapers? Price and stops are universal; the differences are in trip types, locale, carbon data and agent readiness. Multi-city itineraries are supported by Datapika, johnvc, memo23, kaix, scrape.badger and ScrapeBase; automation-lab and Scrape Sage list one-way and round-trip only. Datapika accepts a two-letter ISO code and returns airline and airport names in 30 or more result languages, johnvc lists 29 or more languages and 39 countries, memo23 takes language, currency and market inputs, and automation-lab lists multi-language and multi-currency output. Carbon data varies: Datapika and skootle report grams for the itinerary plus Google's typical figure for the route, automation-lab, scrape.badger and ScrapeBase report kilograms or a percentage versus typical, and johnvc and memo23 document no CO2 field. Skootle adds a price_insights row with median, P10 and P90 fares, a 0 to 100 completeness score and an agentMarkdown field. Kaix keeps the raw Google response next to normalized rows and adds calendar prices. Booking links come from johnvc as a $0.02 per option add-on and from makework36, which compares Google Flights with Kiwi, Travelpayouts, Ryanair, easyJet, Wizz Air and Norwegian. Datapika echoes the query inputs and a scraped_at timestamp on every row so fan-out runs merge by route and date. - Multi-city: Datapika, johnvc, memo23, kaix, scrape.badger, ScrapeBase; not automation-lab or Scrape Sage - Languages: Datapika 30+ via ISO code; johnvc 29+ languages and 39 countries; memo23 language, currency and market - CO2: Datapika and skootle in grams with route typical; automation-lab, scrape.badger, ScrapeBase in kg or percent; johnvc and memo23 none documented - Booking links: johnvc as a paid add-on, makework36 across seven sources; Datapika returns none - Agent extras: skootle agentMarkdown and completeness score; Datapika echoed inputs plus scraped_at on every row ### When should you choose which Google Flights scraper? Pick by the shape of your workload, not the headline rate. If you want the most battle-tested option and can pay about 4 cents per search, johnvc's 2,480 users, 4.0 rating and 100% success on August 29, 2026 make it the low-risk default, and its booking-link add-on is unique among the Google-only actors. If you pull whole result pages in bulk, memo23 at $1.00 per 1,000 plus $0.005 per start with 99.9% success, or automation-lab at $0.23 per 1,000 plus the same start fee, are the cheapest options with real volume behind them. If your agent asks one question at a time and wants an MCP tool that returns the full page, scrape.badger's $0.012 per search is simple to reason about. Choose Datapika at $1.50 per 1,000 when you need CO2 in grams, 30 or more result languages, multi-city trips and a max_results cap of 1 to 200 that bounds itinerary charges before a run starts; accept that platform usage is billed on top and that its 93.5% success trails the leaders. Choose skootle only for its fare percentiles, and skip ScrapeBase until its 0% success rate changes. - Most used, rated, booking links: johnvc, at the highest cost per search - Bulk fare history at the lowest proven cost: memo23 or automation-lab - Agent tool with per-search billing: scrape.badger - CO2 in grams, 30+ languages, multi-city, per-query row cap: Datapika - Fare percentiles and LLM-ready markdown: skootle, at a premium and 59.3% success - Cross-site price comparison with booking links: makework36, despite its 1.0 rating ### Steps 1. Open https://apify.com/openclawai/google-flights-scraper, enter origin and destination IATA codes such as JFK and LAX, a departure date, and optionally a return date or multi-city legs. 2. Set trip type, cabin, passenger mix, a two-letter language code and a max_results cap between 1 and 200; at $0.0015 per itinerary the cap bounds the itinerary charges for that query, with Apify platform usage billed on top. 3. Run it and export JSON, CSV or Excel, or call the run-sync-get-dataset-items endpoint from Python or JavaScript with your Apify token to get the itinerary array in one request. 4. For AI agents, add the actor as a tool through https://mcp.apify.com/?tools=fetch-actor-details,openclawai/google-flights-scraper so a Claude or GPT assistant can request fares directly. ### FAQ Q: Is there an official Google Flights API I should use instead? A: No. Google's QPX Express API, built on the $700 million ITA Software acquisition, was announced for retirement on November 1, 2017 and shut down on April 10, 2018, with Google citing low interest among travel partners. Only the contract-only enterprise product remains. As of August 29, 2026 there is still no public replacement, so a scraper that reads the Google Flights results page is the practical route for independent developers and agents. Q: Is scraping Google Flights legal or allowed? A: The data is public airfare information that any visitor sees without logging in, and none of the actors here require a Google account or bypass access controls. Google's terms of service discourage automated access, so the practical risk is Google throttling or blocking requests, not your credentials. Use the data for research, monitoring and analysis, keep request volumes reasonable, and consult a lawyer if you plan to redistribute fares commercially in your jurisdiction. Q: Why do the prices per 1,000 differ so much between these scrapers? A: Because they bill different units. Memo23, automation-lab, skootle and Datapika charge per row, and all but Datapika add a fee per run start while Datapika adds Apify platform usage instead. Johnvc charges $0.03 per results page plus $0.01 per run, scrape.badger $0.012 per search regardless of rows, kaix $0.06 per 1,000 plus compute, and ScrapeBase $19.99 a month. Convert each to your actual rows per query before comparing; the worked example above does that. Q: Which Google Flights scraper has the best success rate? A: On Apify's 30-day metric as of August 29, 2026, automation-lab, johnvc, makework36, leorochasantos and Scrape Sage show 100%, memo23 99.9%, scrape.badger 98.2%, Datapika 93.5%, kaix 92.8%, skootle 59.3% and ScrapeBase 0%. Weight that by volume: johnvc's figure covers roughly 29,800 runs in 30 days, leorochasantos 30 and Scrape Sage 51. A failed Datapika run pushes no rows and bills no itinerary events, and each search retries three times with backoff before giving up. Q: Can I get Google Flights data through MCP for an AI agent? A: Yes. Every Apify actor can be exposed as a tool through Apify's MCP server, and both scrape.badger and Datapika document that path. For Datapika the tool URL is https://mcp.apify.com/?tools=fetch-actor-details,openclawai/google-flights-scraper. The agent sends airports, date, trip type and cabin and receives typed rows with price, airline IATA codes, segments, CO2 in grams and the echoed query inputs, so fan-out across dates can be merged without parsing HTML. Q: How many itineraries does a single Google Flights search return? A: Google usually renders 30 to 80 itineraries for one route and date. Datapika returns up to your max_results, default 25 and maximum 200, and bills itinerary events only for rows delivered, so a typical search runs between 4.5 and 12 cents before platform usage. Automation-lab also caps at 200 per search. Per-search actors such as scrape.badger return the full page for one $0.012 charge, which is why that model wins on cost only when you want every result. ## LinkedIn has no public jobs API: 6 working alternatives compared Canonical: https://datapika.com/compare/linkedin-jobs-api-alternative Updated: 2026-08-29 LinkedIn has no public jobs API: its Job Posting API is closed to new partners and only lets approved ATS vendors post jobs, not read them (Microsoft Learn, updated June 3, 2026). The working alternatives are Apify actors that read public job pages without a login. As of August 29, 2026, valig is cheapest at $0.28 per 1,000 jobs, curious_coder has the largest base at 142,306 users for $1.00 per 1,000, and Datapika's job-board-scraper costs $5 per 1,000 but sweeps LinkedIn plus 7 other boards in one deduplicated run, targets employers by LinkedIn company ID, and holds a 5.0 rating. ### Is there a LinkedIn Jobs API for third parties? No, and there never has been a self-serve one. LinkedIn's developer docs (Microsoft Learn, updated June 3, 2026) split access into three tiers. Open Permissions, the only ones any developer can enable, cover Sign In with LinkedIn and Share on LinkedIn. Marketing, Sales Navigator, and Talent programs all need explicit approval. Job data sits under Talent Solutions: Recruiter System Connect, Apply Connect, Apply with LinkedIn, and Premium Job Posting, each gated behind a partner application form. The Job Posting API overview page carries a notice that reads: 'We are currently not accepting new partnerships for LinkedIn's Job Posting API.' Even for approved partners, that API pushes jobs into LinkedIn on behalf of employers. It has no endpoint for searching or reading listings, so the phrase 'LinkedIn jobs API' describes a product that was never sold to the public. Every practical option reads the public jobs pages instead. - Self-serve permissions as of June 2026: profile, email, and w_member_social. No jobs scope exists. - Talent Solutions APIs require a partner application, an API agreement with data restrictions, and a LinkedIn relationship manager, per the Job Posting legal requirements section. - The Job Posting API is write-only: it syncs Basic or Promoted jobs into LinkedIn and returns nothing you could call a search result. - No Talent Solutions API has a published price; quotes arrive only after partnership approval. ### Which LinkedIn jobs scraper should you use instead? Six Apify actors cover most of the market, and every number below comes from each actor's live store page on August 29, 2026. curious_coder/linkedin-jobs-scraper is the incumbent: 142,306 users, 14,458 of them in the last 30 days, a 4.56 rating, from $1.00 per 1,000 results, and 99.7% of runs succeeding. cheap_scraper's 'Remove Duplicate Jobs' actor has 47,084 users and a 4.02 rating at $0.35 to $0.70 per 1,000 depending on your Apify plan, with a documented cap of 1,000 jobs per search. bebity charges $29.99 per month plus usage, rates 4.28 with 35,315 users, but only 257 people used it in the last 30 days, and Apify retires rental pricing on October 1, 2026. valig is the price floor at $0.28 per 1,000 with a 4.63 rating and 100% run success. apimaestro's 'No Cookies' actor costs $5.00 per 1,000 and rates 3.73. Datapika's job-board-scraper also bills $5 per 1,000 (2,471 users, 381 in the last 30 days, 5.0 from 3 reviews, 100% run success) and is the only row covering 8 boards with cross-board deduplication. Datapika loses on per-row price; it wins on rating, source coverage, and agent access. - Cheapest per row: valig at $0.28 per 1,000, then cheap_scraper at $0.35 to $0.70. - Largest install base: curious_coder at 142,306 users, roughly 3x cheap_scraper and 57x Datapika. - Highest rating: Datapika at 5.0 (3 reviews), then valig at 4.63 and curious_coder at 4.56. - Multi-board sweep, LinkedIn company-ID targeting, and an MCP endpoint: Datapika only. - Explicit no-login, no-cookie operation stated on the listing: cheap_scraper, apimaestro, and Datapika. ### What does it cost to scrape 1,000 or 10,000 LinkedIn jobs? Pay-per-result actors make the arithmetic simple: multiply the row count by the per-row rate. Rental actors do not, because a flat monthly fee is cheap at volume and expensive at low volume. Using the August 29, 2026 store prices: valig at $0.00028 per job costs $0.28 for 1,000 jobs and $2.80 for 10,000. curious_coder at $0.001 per job costs $1.00 and $10.00. Datapika at $0.005 per job costs $5.00 and $50.00, so on a LinkedIn-only pull it is 5x curious_coder and about 18x valig. bebity's $29.99 rental works out to $29.99 per 1,000 if you only need 1,000 jobs a month, or $3.00 per 1,000 at 10,000 jobs, before Apify compute charges on top. The Datapika premium buys deduplication: a job posted on LinkedIn, Indeed, and Glassdoor is billed once as a single row tagged with all three sources, so a 10,000-row multi-board pull costs $50 rather than three separate actor bills. If you only want LinkedIn rows at volume, the same catalog has a $0.0005 per job LinkedIn-only option, which prices 10,000 jobs at $5.00. - 1,000 jobs: valig $0.28, curious_coder $1.00, cheap_scraper $0.35 to $0.70, apimaestro $5.00, Datapika $5.00, bebity $29.99 plus usage. - 10,000 jobs: valig $2.80, curious_coder $10.00, cheap_scraper $3.50 to $7.00, apimaestro $50.00, Datapika $50.00, bebity $29.99 plus usage. - Break-even for bebity against curious_coder is about 30,000 jobs per month ($29.99 divided by $0.001), and that comparison ends when rental billing stops on October 1, 2026. - Datapika's $5 covers up to 8 boards per row; the LinkedIn-only budget option in the same catalog is $0.0005 per job. ### Can you scrape LinkedIn jobs without login or cookies, and what fields do you get? Yes. LinkedIn serves a public version of job search to logged-out visitors, which is why curious_coder's README tells you to copy search URLs from an incognito window. cheap_scraper, apimaestro, and Datapika state on their listings that no credentials are needed; apimaestro's page frames it as not risking your account by sharing cookies. Cookie-based scrapers exist for the logged-in feed, but they expose your own profile to restriction, which is why 'No Cookies' and 'No Login' now appear in actor titles. The trade-off is scope. cheap_scraper documents a 1,000-result ceiling per search, Datapika's listing notes LinkedIn rate-limits to about 100 results per IP and recommends residential proxies for large runs, and curious_coder's page reports that LinkedIn's August 2026 AI-powered search honors only four URL filters: date posted, company, easy apply, and under 10 applicants. Datapika's schema is the widest of the group because it normalizes LinkedIn rows against 7 other boards. - Core fields on every Datapika row: title, company, location, job URL, direct apply URL, remote flag, and job type. - Salary: min, max, currency, and interval, with optional annual normalization for cross-board comparison. - Company enrichment: size, revenue, rating, review count, industry, and logo. - Full description in markdown or HTML when linkedinFetchDescription is on. - Targeting: linkedinCompanyIds restricts a run to named employers; hoursOld filters recency but cannot be combined with easyApply on LinkedIn, per the store listing. - curious_coder adds job-poster details; cheap_scraper adds a saveOnlyUniqueItems switch to avoid paying for duplicates within LinkedIn. ### When to choose which LinkedIn jobs API alternative Pick by the shape of the job, not by the headline price. A one-off LinkedIn keyword dump has a different cost curve from a weekly competitor watch across every board an employer posts on, and an AI agent that needs a tool endpoint has different requirements again. Volume matters too: below roughly 30,000 jobs a month every pay-per-result actor in the table undercuts bebity's $29.99 rental, and above that the rental only wins until Apify retires the model on October 1, 2026. Check run success before price, because a cheap actor that fails on the URLs you feed it costs engineering hours instead of cents: valig, apimaestro, and Datapika each report 100% run success, curious_coder and cheap_scraper 99.7%, and bebity 99.2%. All figures are as of August 29, 2026. - Lowest cost, LinkedIn only: valig ($0.28 per 1,000, 4.63) or cheap_scraper ($0.35 to $0.70, built-in dedup, 1,000-result cap per search). - Safest default for a large team: curious_coder ($1.00 per 1,000, 142,306 users, 99.7% run success), accepting that LinkedIn's new search limits URL filters. - Tracking named employers across boards: Datapika ($5 per 1,000, LinkedIn company-ID targeting, one deduplicated row per job across 8 boards, 5.0 rating). - AI agents and MCP clients: Datapika, reachable through the Apify MCP server with pay-per-result billing and no API key of its own. - High-volume LinkedIn-only feeds: the $0.0005 per job LinkedIn-only option in the Datapika catalog. - Avoid starting a new pipeline on bebity's $29.99 rental; Apify retires rental pricing on October 1, 2026 and migrates remaining rental actors to pay-per-usage billing. ### Steps 1. Open https://apify.com/openclawai/job-board-scraper, enter up to 5 search terms and a location, and select LinkedIn plus any of the other 7 boards you want deduplicated into the same dataset. 2. Add linkedinCompanyIds to restrict the run to named employers, switch on linkedinFetchDescription for full posting text and direct apply URLs, and set hoursOld for recency (do not pair it with easyApply on LinkedIn). 3. Start the run; LinkedIn rows typically land within 5 to 20 seconds, and you pay $0.005 per delivered job. Export JSON or CSV from the dataset view or pull it through the Apify API. 4. For agents, register the MCP server at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper so the scraper appears as a callable tool with the same per-result billing. ### FAQ Q: Does LinkedIn have an official jobs API? A: Not for reading jobs. LinkedIn's Job Posting API, documented on Microsoft Learn (updated June 3, 2026), is limited to approved Talent Solutions partners such as ATS vendors and job distributors, and it only posts jobs into LinkedIn. The page states LinkedIn is not accepting new partnerships for it. Self-serve Open Permissions cover sign-in and sharing only. There is no endpoint that returns job search results to a third party. Q: Is it legal or allowed to scrape LinkedIn jobs without an API? A: In the US, the Ninth Circuit held on April 18, 2022 (hiQ Labs v. LinkedIn) that scraping publicly accessible pages that need no account does not violate the Computer Fraud and Abuse Act. The same case ended on December 6, 2022 with a consent judgment: a $500,000 award to LinkedIn, a permanent injunction, and hiQ's stipulation to liability for trespass and misappropriation after the court found it breached the User Agreement. Logged-out scraping of public job pages avoids that agreement; check local law and your own use of the data. Q: Can I scrape LinkedIn jobs without logging in or sharing cookies? A: Yes. LinkedIn exposes a public job search to logged-out visitors, and cheap_scraper, apimaestro, and Datapika all state on their Apify listings (as of August 29, 2026) that no credentials are required. The limits are LinkedIn's: cheap_scraper documents 1,000 results per search, Datapika notes about 100 results per IP and recommends residential proxies for large runs, and only four URL filters survive LinkedIn's August 2026 search redesign. Q: Why pay Datapika $5 per 1,000 when valig charges $0.28? A: Because the rows are not the same product. valig returns LinkedIn listings for LinkedIn searches at the lowest price in the table. Datapika returns one deduplicated row per job across LinkedIn, Indeed, Glassdoor, Google Jobs, ZipRecruiter, Naukri, Bayt, and BDJobs, with salary normalization, company enrichment, and LinkedIn company-ID targeting. For a LinkedIn-only feed, use valig or the $0.0005 per job LinkedIn-only option in the Datapika catalog. Q: What happens to bebity's $29.99 per month LinkedIn Jobs Scraper? A: Apify's rental pricing documentation states that no new rental actors could be published or re-priced after April 1, 2026, and that on October 1, 2026 rental actors are fully retired and migrated to pay-per-usage pricing. bebity's listing still showed $29.99 per month plus usage on August 29, 2026, with 257 monthly users, so expect its price to change and check the pricing tab before building a recurring job on it. Q: Can an AI agent call a LinkedIn jobs scraper through MCP? A: Yes. Datapika's job-board-scraper is exposed through the Apify MCP server at mcp.apify.com with the fetch-actor-details tool, so a Claude, ChatGPT, or custom agent can discover the input schema, run a search, and read the dataset without a LinkedIn account or a separate API key. Billing stays at $0.005 per job, the same rate a human pays in the Apify console, which keeps agent spend predictable. ## Google Places API alternatives: contact fields at a flat $3.00 per 1,000 places Canonical: https://datapika.com/compare/google-maps-api-alternative Updated: 2026-08-29 For bulk place records with phone, website, hours, and rating, Apify actors undercut the official API for most jobs as of August 29, 2026. Google Places API (New) bills $20 per 1,000 Place Details Enterprise requests and forbids storing anything except place IDs. Datapika's Google Maps scraper costs a flat $0.003 per place at every volume ($3.00 per 1,000, $30.00 per 10,000), needs no API key, and runs over MCP. compass/crawler-google-places (580,768 users, 4.69 rating) has the lowest headline rate, from $1.50 per 1,000 plus paid add-ons. Outscraper gives 500 places free, then $3 per 1,000. Use Google for live lookups inside a Google Map, an actor for stored datasets. ### Why is the Google Places API so expensive for bulk place data? Google's Places API (New) pricing page, last updated August 25, 2026, bills per request and by SKU, not per place. Text Search costs $32 per 1,000 requests on Essentials and Pro and $35 on Enterprise. Place Details costs $5 per 1,000 on Essentials, $17 on Pro, $20 on Enterprise, and $25 on Enterprise + Atmosphere. The SKU is decided by the fields you request. Per the Place Details field documentation, formattedAddress and location are Essentials, displayName is Pro, and nationalPhoneNumber, websiteUri, regularOpeningHours, rating, userRatingCount, and priceLevel all trigger Enterprise. Reviews sit in Enterprise + Atmosphere. Contact data therefore costs the $20 tier. Free monthly allowances are 10,000 Essentials, 5,000 Pro, and 1,000 Enterprise requests per SKU. Two structural limits matter more than the rate card. Text Search (New) returns at most 20 results per page and 60 per query, so a city-wide category needs dozens of sub-queries. And the Places API policies say you must not pre-fetch, cache, or store Places API content, with only the place ID exempt, so a stored spreadsheet of phones and hours sits outside the terms. Volume discounts start late: Text Search Pro drops to $25.60 only above 100,000 monthly requests. - Place Details Enterprise (phone, website, hours, rating, price level): $20 per 1,000 requests, 1,000 free per month (Google pricing page, updated August 25, 2026) - Text Search Essentials and Pro $32, Enterprise $35 per 1,000 requests; max 20 results per page, 60 per query - Caching policy: only place_id may be stored; every other field must be re-fetched - Needs a Google Cloud project, billing account, and API key before the first call ### Google Places API alternatives compared: price per 1,000, users, ratings Prices are per 1,000 delivered places, except Google (per 1,000 requests) and SerpApi (per 1,000 searches, about 20 places each), all checked August 29, 2026. compass/crawler-google-places is the incumbent: 580,768 users, 34,835 in the last month, a 4.69 rating, and 97.9 percent run success. It advertises from $1.50 per 1,000 places, but lists separate charges for additional details, contact enrichment, reviews (per review) and social profiles, and the base rate falls as your Apify plan rises. Sibling compass/google-maps-extractor (98,106 users, 4.89 rating) charges $5.00 per 1,000 on the Free plan, $2.10 on Business. Datapika bills one event per place at a flat $0.003 at every volume, plus a $0.00005 actor-start event that never adds up to a cent, so $3.00 per 1,000 and $30.00 per 10,000. The Apify API reports 559 completed runs as of August 29, 2026, so it trails compass on community size. It does not beat compass's from-$1.50 headline rate. Its case is a single rate that does not move with your Apify plan, your volume, or the fields you request, no key or login, and an MCP tool with a four-field input. Outscraper charges $0 for the first 500 places each month, $3 per 1,000 up to 100,000, then $1. SerpApi sells searches: $25 per month for 1,000, $75 for 5,000, with a 250-search free plan. - Largest user base: compass/crawler-google-places, 580,768 users and 4.69 rating (Apify Store, August 29, 2026) - Lowest headline rate: compass from $1.50 per 1,000 before add-ons; flattest rate: Datapika at $0.003 per place at every volume, no add-ons - Cheapest first test: Outscraper's 500 free places, or Google's 1,000 free Enterprise requests - No API key, no login, MCP-ready: Datapika, compass, Outscraper; SerpApi and Google need a key ### What do 1,000 and 10,000 places with contact data actually cost? Assume you need name, address, phone, website, hours, and rating for every place. List prices, checked August 29, 2026. Google Places API (New), using Text Search Pro to find place IDs and one Place Details Enterprise request per place: 1,000 places need at least 51 search requests (free under the 5,000 Pro cap) plus 1,000 details requests at $20 per 1,000, so $20.00, or $0 while the 1,000 free Enterprise requests last. At 10,000 places it is 9,000 paid details requests after the free 1,000, so $180.00 per month. There is a cheaper Google path. Requesting Enterprise fields inside Text Search costs $35 per 1,000 requests, and 10,000 places need at least 501 requests (167 queries capped at 60 results, three pages each), about $17.54, with 1,000 requests free each month. On paper that undercuts Datapika. In practice you must write 167 non-overlapping queries, deduplicate heavily, and still cannot store the output. Datapika: 1,000 places = 1,000 x $0.003 = $3.00. 10,000 places = 10,000 x $0.003 = $30.00. 100,000 places = $300. The $0.00005 actor-start event adds well under a cent per run. compass at the from-$1.50 rate: $1.50 and $15.00 before add-ons. Outscraper: 500 free + 500 x $0.003 = $1.50; 9,500 x $0.003 = $28.50. - 1,000 places: Google Place Details Enterprise $20.00; Datapika $3.00; compass from $1.50; Outscraper $1.50 - 10,000 places: Google $180.00 (after 1,000 free); Datapika $30.00; compass from $15.00; Outscraper $28.50 - Google Text Search Enterprise: about $17.54 per 10,000 places, but 167 queries of 60 results and no storage allowed - SerpApi: 10,000 places is at least 500 searches, half of the $25 per month Starter plan's 1,000 searches ### Which fields do you get from each option? The official API returns the freshest record but meters every field. Phone, website, opening hours, rating, rating count, and price level cost the Enterprise SKU; reviews, editorial summaries, and amenity flags cost Enterprise + Atmosphere. You also get stable place IDs, which you may store indefinitely. Datapika returns a flat JSON record per place with title, category, address, phone, website, review_rating, review_count, status, description, price_range, open_hours keyed by weekday, latitude, longitude, plus_code, timezone, up to 5 image URLs, Google's cid, and a direct Maps link. Switching scrapeReviews on nests the top 10 reviews inside the record, adding about 3 seconds per place and no per-review charge. It does not return email addresses or reviewer profile links. compass/crawler-google-places goes wider when you pay for it: reviews with reviewer details, images, and contact info including full name, email, and job title, with contacts, social profiles, and email verification priced as add-ons. It also splits a large area into subregions internally to get past Google's 120-places-per-area limit. SerpApi returns Google's local results with phone, rating, and operating hours, 20 per page. Outscraper matches the core contact set and sells reviews separately at $3 per 1,000. - Google: every field metered; phone, website, hours, rating, price level = Enterprise; reviews = Enterprise + Atmosphere - Datapika: 18 core fields plus top 10 reviews per place, no emails, no reviewer profiles - compass: emails, job titles, social profiles, and per-review billing as paid add-ons - Every scraper faces the roughly 120-place ceiling per Maps search; compass automates the subregion split, Datapika leaves it to you ### When should you use the Places API, compass, or Datapika? Choose the official Places API when the data is displayed on a Google Map inside your product, for a handful of live lookups per session, or when you must stay strictly inside Google's caching and attribution terms. Under 1,000 Enterprise requests a month it is free, and it is the only option with a supported, versioned contract. Choose compass/crawler-google-places when you need the widest enrichment (emails, social profiles, full review histories), the largest track record (580,768 users, 97.9 percent success as of August 29, 2026), the lowest headline rate at from $1.50 per 1,000, or automatic subregion splitting so one area yields far more than 120 places. Choose Datapika when the job is a clean per-place dataset at a predictable price: a flat $0.003 per place, so $3.00 per 1,000, $30.00 per 10,000, $300 per 100,000, billed only for places delivered, with no Google Cloud project, no key, and no subscription. It is also the simplest option for an AI agent, because the MCP tool takes a query string and returns records without glue code. Choose Outscraper for one-off jobs under 500 places, and SerpApi if you already pay for it. Whatever you pick, plan around the roughly 120-result ceiling per Maps search and deduplicate on cid or link when merging. - Official API: in-product Google Map display, low volume, strict compliance - compass: maximum enrichment, lowest headline rate and largest user base, more expensive with add-ons - Datapika: one flat $0.003 per place at any volume, no add-ons, no key, MCP-native - Every scraper: split broad queries by neighborhood or category to beat the 120-place ceiling ### Steps 1. Open https://apify.com/openclawai/google-maps-scraper with a free Apify account, or register it as an agent tool at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/google-maps-scraper. No Google Cloud project, billing account, or API key is needed. 2. Replace each Text Search call with one query string per area (for example, dentists in Camden London), keep maxResults at or below about 120 per query, and set scrapeReviews to true if you need review text. 3. Run from the console, the REST run-sync-get-dataset-items endpoint, the CLI, or the MCP tool, then map title, phone, website, open_hours, and review_rating onto the Place Details field names your code already uses. 4. Deduplicate on cid or link when merging neighborhood queries, schedule recurring runs for fresh snapshots, and pay a flat $0.003 per delivered place at any volume. ### FAQ Q: Why not just use the official Google Places API? A: For low volumes displayed inside a Google Map it is the right choice. The problem is price and terms at scale. Per Google's pricing page dated August 25, 2026, a Place Details request that includes phone, website, hours, and rating bills at the Enterprise SKU, $20 per 1,000, after 1,000 free requests a month. Text Search caps at 60 results per query, and the policies prohibit storing any field except place_id. A 10,000-place contact dataset costs about $180 a month there versus $30.00 on Datapika. Q: Is scraping Google Maps legal or allowed? A: Collecting publicly visible business listings is generally treated as lawful in the US and EU when the data is public, non-personal business information and you respect rate limits. Google's Terms of Service bind accounts that agree to them, which is why scrapers do not log in. Reviews contain reviewer names, which are personal data under GDPR, so store them only with a lawful basis. This is not legal advice; check your jurisdiction and intended use before building a product on scraped data. Q: How many places can one Google Maps query return? A: About 120 for any single search, because that is roughly how many results Google Maps renders no matter how far you scroll. Datapika's input accepts maxResults up to 500, but a broad query still stops near 120. The official Text Search (New) API is stricter at 60 results per query across three pages of 20. To collect thousands of places, run one query per neighborhood, postcode, or sub-category and merge the datasets, dropping duplicates on the cid or link field. Q: How does Datapika's per-place pricing compare with compass/crawler-google-places? A: compass advertises from $1.50 per 1,000 places and has 580,768 users and a 4.69 rating as of August 29, 2026, but its store page bills details, contact enrichment, reviews (per review), and social profiles as separate events, and the base rate depends on your Apify plan. Datapika charges one event per place at a flat $0.003, so 1,000 places is $3.00 and 10,000 is $30.00, with the top 10 reviews per place included when enabled. On headline rate compass is cheaper; Datapika's advantage is that the number does not move with your plan, your volume, or the fields you request. Q: Can an AI agent call a Google Maps scraper without a Google API key? A: Yes. Datapika's actor is exposed through the Apify MCP server at mcp.apify.com with openclawai/google-maps-scraper in the tools list, so Claude, Cursor, or any MCP client can read the input schema, run a query, and consume the place records. The only credential is the Apify token, which also carries the pay-per-result billing. There is no Google Cloud project, no OAuth, and no quota to configure, and an agent can compute cost before calling because each place is a fixed $0.003 event. Q: Does Datapika return email addresses or full review histories? A: No. The record covers business-level data: name, category, address, phone, website, rating, review count, status, hours, coordinates, plus code, timezone, images, and a Maps link, plus the top 10 reviews per place when scrapeReviews is on. There are no email addresses, reviewer profile links, or complete review histories. If you need emails or social profiles, compass/crawler-google-places sells those as paid add-ons, or you can run an email-enrichment step on the website field. ## Best Indeed scrapers: $0.10 to $6.00 per 1,000 jobs, compared Canonical: https://datapika.com/compare/best-indeed-scrapers Updated: 2026-08-29 For cheap Indeed-only pulls, valig/indeed-jobs-scraper is the best value at $0.10 per 1,000 jobs on the free plan, with a 5.0 rating from 17 reviews and 26,654 users (Apify Store, August 29, 2026). The category leader, misceres/indeed-scraper, has the most users at 29,854 but bills $6.00 per 1,000 on the free plan and rates 3.76 from 64 reviews. Datapika's job board scraper is the pick when Indeed is one of several sources: $0.005 per job ($5 per 1,000), 5.0 rating, 2,471 users, and one run sweeps Indeed with LinkedIn, Glassdoor, ZipRecruiter, Naukri, and Bayt, deduplicated, with salary and company fields. ### Why not use the official Indeed API? Because there is no self-serve Indeed API for reading job postings. Indeed rolled back the Publisher Program to new applicants in October 2022 and retired the Publisher API in 2023 without a self-serve replacement (Job Boardly, Indeed affiliate program status, 2026; JobsPipe, Indeed API key explainer). The misceres scraper's own README, live on August 29, 2026, states that the Publisher Jobs API (Get Job and Job Search) has been deprecated. What Indeed still offers is employer-side. The Job Sync API publishes and updates an employer's own postings, Indeed Apply receives applications, and Sponsored Jobs manages campaigns, per docs.indeed.com. Job Sync can return a posting's status and details, but only to the ATS that created it under an OAuth token. None of these let a third party search or read other employers' listings, and access is a partner review rather than a signup form. Enterprise data licensing exists but is NDA-gated with no published pricing. That leaves scrapers as the only practical way for a job board, researcher, or AI agent to read Indeed postings. The Apify Store lists dozens of Indeed actors; this page ranks the six most visible Indeed-only ones plus Datapika. - October 2022: Publisher Program rolled back to new applicants and still closed in 2026 (Job Boardly) - 2023: Publisher API retired, existing keys stopped working, no self-serve equivalent since (JobsPipe) - Remaining Indeed APIs move data toward Indeed, not out of it: Job Sync, Indeed Apply, Sponsored Jobs (docs.indeed.com) - Data licensing is enterprise-only and NDA-gated with no public price - Every actor below reads public listings without an Indeed account or API key ### How the 7 Indeed scrapers compare on price, rating, and users The table is a snapshot taken August 29, 2026 from the store pages and the public Apify store API. Three facts stand out. First, the advertised "from" price is often a Gold-plan price: misceres shows from $3.00 per 1,000 but bills $6.00 on the free plan, and automation-lab shows from $1.80 but bills $3.00 per 1,000 plus $0.005 per run below Gold. Second, price and popularity are inversely related: misceres leads on users (29,854) with a 3.76 rating from 64 reviews, while valig charges $0.10 per 1,000, 60 times less at base tier, and holds 5.0 from 17 reviews with 26,654 users. Third, borderline costs $5.00 flat and keeps a 4.63 rating from 37 reviews. Datapika bills $5.00 per 1,000 and does not win on price. On raw 30-day success it is last among the working actors at 90%, because aborted and timed-out runs count against it. Its row wins on scope: one input covers Indeed, LinkedIn, Glassdoor, ZipRecruiter, Naukri, and Bayt by default (Google Jobs and BDJobs are selectable but currently return nothing), deduplicates across them, and returns salary and company fields where the board exposes them. Its README documents MCP access for agents. - Cheapest: valig at $0.10 per 1,000 on the free plan, $0.07 on Gold and above, 99.7% raw success over the last 30 days - Most used: misceres at 29,854 users and 2,299 monthly, but the lowest rating in the set at 3.76 - Highest-rated with real volume: valig (5.0, 17 reviews) and borderline (4.63, 37 reviews) - Most expensive at base tier: misceres at $6.00 per 1,000; borderline and Datapika are $5.00 flat - Avoid for now: smorgi_apps, offline until August 30, 2026 and then rising to about $15.00 per 1,000 per its README ### What 1,000 and 10,000 Indeed jobs actually cost All seven actors bill per delivered row, so the arithmetic is linear. Using free-plan prices from the store pages and API on August 29, 2026: valig at $0.10 per 1,000 means 1,000 jobs cost $0.10 and 10,000 cost $1.00, plus $0.001 per run per GB of memory. Misceres at $6.00 per 1,000 means $6.00 for 1,000 and $60.00 for 10,000; a Gold plan drops that to $3.00 and $30.00. Automation-lab at $3.00 per 1,000 plus $0.005 per run means $3.005 for 1,000 in one run and $30.005 for 10,000. Factden at $2.00 per 1,000 means $2.00 and $20.00. Borderline and Datapika at $5.00 per 1,000 mean $5.00 and $50.00. The Datapika number needs one adjustment. Its rows are deduplicated across boards, so a posting on Indeed, LinkedIn, and Glassdoor is delivered and billed once, tagged with every source. Buying 10,000 rows from an Indeed scraper and 10,000 from a LinkedIn scraper means paying twice for each cross-posted job and merging the files yourself. For an Indeed-only feed, valig's $1.00 per 10,000 is unbeatable. For a multi-board dataset with normalized salaries and company size, revenue, rating, and industry, Datapika's $50 per 10,000 replaces several scrapes and a dedup step. - valig: $0.10 for 1,000 jobs, $1.00 for 10,000, plus $0.001 per run per GB - misceres: $6.00 for 1,000, $60.00 for 10,000 on the free plan; $3.00 and $30.00 on Gold - automation-lab: $3.005 for 1,000 in a single run, $30.005 for 10,000; $1.80 per 1,000 on Gold - factden: $2.00 for 1,000, $20.00 for 10,000; $1.20 per 1,000 on Gold and above - borderline and Datapika: $5.00 for 1,000, $50.00 for 10,000; Datapika rows are deduplicated across boards - New Apify accounts get $5 in free monthly credit, which factden's page notes is roughly 2,500 jobs at $2.00 per 1,000 ### Field coverage: salary, company data, and descriptions Price per 1,000 is only half the comparison; the rows differ in shape. Datapika returns on each row the title, company, location, job URL, a direct apply URL when available, and the full description in markdown or HTML, plus salary minimum, maximum, currency, and interval with an enforceAnnualSalary option that converts wages to yearly equivalents, plus company size, revenue, rating, industry, and logo where the board exposes them. Filters cover remote only, job type, posted within N hours, distance radius, and easy apply, and a country setting switches Indeed between markets such as the US, UK, or India. Among the Indeed-only actors, enrichment varies. Factden's page highlights a native salary-range filter, every result page past the roughly 1,000-result ceiling, and free company profiles (rating, size, industry, CEO, founding year, website) in a separate dataset, with no login. Automation-lab's page states no Indeed account is needed. Misceres remains the reference for a plain listing schema; valig emphasizes price and reliability over field depth. Check whether pay comes back as structured fields or a raw string; Datapika does the former and normalizes hourly, monthly, and annual figures so a $45 per hour contract sorts next to a $95,000 salaried role. - Datapika: structured salary (min, max, currency, interval) with optional annual normalization - Datapika: company size, revenue, rating, industry, logo, and description alongside each posting where available - Datapika: full description in markdown or HTML, direct apply URL when available, remote flag, job type - factden: salary-range filter, all result pages, separate free company-profiles dataset, no login (store page, August 29, 2026) - automation-lab: no Indeed account or API key required; pricing revised August 25, 2026 to $3.00 per 1,000 plus $0.005 per run on the free plan - All seven actors read public listings only; none returns applicant or employer private data ### Which Indeed scraper should you choose? Match the tool to the job. If you need raw Indeed listings at the lowest cost and will enrich them yourself, choose valig: $0.10 per 1,000, a 5.0 rating, and 99.7% of 214,408 runs succeeding in the last 30 days (August 29, 2026). If you want the longest track record and largest community, misceres has 29,854 users, but budget $6.00 per 1,000 on the free plan ($3.00 on Gold) and read the 64 reviews behind its 3.76 rating first. Borderline is the middle ground for teams that value a 4.63 rating from 37 reviews over price; its 96.5% raw success rate is the lowest of the three big actors. Choose Datapika when Indeed is one input among several. Salary research, hiring-signal monitoring, and feeds that must not double-count cross-posted roles benefit from one deduplicated dataset instead of separate Indeed, LinkedIn, and Glassdoor exports, and its README documents MCP access for AI agents. Its 90% raw success rate counts user-aborted and timed-out runs. Skip smorgi_apps, offline until August 30, 2026 with a rise to about $15 per 1,000 announced in its README. Factden and automation-lab are promising but lightly tested: near-perfect rates on only 996 and 873 runs. - Cheapest Indeed-only feed: valig at $0.10 per 1,000 - Largest user base and longest history: misceres at $6.00 per 1,000 on the free plan, 3.76 rating - Best reviews among the $5-per-1,000 Indeed-only actors: borderline at 4.63 from 37 reviews - Multi-board dedup, salary normalization, company enrichment, and documented MCP access: Datapika at $5.00 per 1,000 - Salary-range filtering on Indeed alone: factden at $2.00 per 1,000 - Not yet: smorgi_apps, offline until August 30, 2026 and moving to about $15 per 1,000 ### Steps 1. Open the Datapika job board scraper at https://apify.com/openclawai/job-board-scraper, enter your search terms and a location, and select Indeed plus any of LinkedIn, Glassdoor, ZipRecruiter, Naukri, or Bayt for the same sweep. 2. Set the Indeed country, turn on the enforceAnnualSalary option if you need comparable pay figures, and add filters such as remote only, job type, or posted within N hours. 3. Run it and export the deduplicated dataset as JSON or CSV, or pull it through the Apify API; you are billed $0.005 per delivered job with no subscription. 4. For AI agents, connect the Apify MCP server at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/job-board-scraper (OAuth sign-in on first connect) so the agent can call the scraper as a tool and pay per result. ### FAQ Q: Does Indeed have an official API for job listings? A: Not for reading postings. Indeed rolled back its Publisher Program to new applicants in October 2022 and retired the Publisher API in 2023 without a self-serve replacement, per Job Boardly and JobsPipe. The APIs that remain, Job Sync, Indeed Apply, and Sponsored Jobs, are employer-side and issued to approved partners after a review, per docs.indeed.com. Enterprise data licensing exists but is NDA-gated. As of August 29, 2026, scrapers are the only practical way to read Indeed listings. Q: Which Indeed scraper is the cheapest? A: valig/indeed-jobs-scraper at $0.10 per 1,000 jobs on the free plan, dropping to $0.07 per 1,000 on Gold and above, per the Apify store API on August 29, 2026. It is also well rated at 5.0 from 17 reviews with 26,654 users, and 99.7% of its 214,408 runs in the last 30 days succeeded. For 10,000 raw Indeed listings that comes to $1.00, compared with $60.00 on misceres at its free-plan price. Q: Why is the most popular Indeed scraper rated only 3.76? A: misceres/indeed-scraper has 29,854 users, the most of any Indeed actor, but its 3.76 rating comes from 64 reviews, more than any competitor, and it bills $6.00 per 1,000 on the free plan, 60 times valig's price; the advertised $3.00 needs a Gold plan. Its 99.1% raw success rate over 175,968 runs in the last 30 days is strong, so the rating reflects price and feature expectations more than failures. User count on Apify mostly tracks how long an actor has been listed. Q: When is Datapika the better choice over an Indeed-only scraper? A: When Indeed is one of several boards you need. Datapika runs Indeed, LinkedIn, Glassdoor, ZipRecruiter, Naukri, and Bayt in one input (Google Jobs and BDJobs are selectable but currently return no results), merges cross-posted jobs into one row tagged with every source, and returns structured salary fields plus company size, revenue, rating, and industry where available. At $0.005 per unique job it costs 50 times valig per row but replaces multiple scrapes and a dedup step. Its README documents MCP access for AI agents. Q: Is scraping Indeed legal or allowed? A: Every actor on this page reads publicly visible listings without logging in, the same data any visitor sees. US courts have generally treated collection of public web pages as outside computer-fraud law, but Indeed's terms of service prohibit automated access, so you carry contractual risk if you hold an Indeed account. Do not collect applicant or employer private data, keep request rates modest, and consult counsel for your jurisdiction and use case, especially in the EU where personal data rules apply to named recruiters. Q: How reliable are the newer Indeed scrapers with near-perfect success rates? A: Treat them as promising rather than proven. factden (44 users, 25 monthly) and automation-lab (211 users, 36 monthly) show 99.8% and 99.7% raw success on August 29, 2026, but on 996 and 873 runs in the last 30 days per the Apify store API, versus 214,408 runs for valig and 175,968 for misceres. A near-perfect rate over a few hundred runs says less than 99.1% over 175,000. Test them on a small batch before moving a production feed. ## Reddit API alternatives that still work after the May 2026 shutdown Canonical: https://datapika.com/compare/reddit-api-alternative Updated: 2026-08-29 The best Reddit API alternative in 2026 is a pay-per-result scraper that reads public pages, because Reddit closed self-serve API registration on November 11, 2025 and began returning 403 on anonymous .json requests in the last week of May 2026. For threads with comments, Datapika's Reddit actor is cheapest at $3 per 1,000 posts including every comment. For bare post listings, Automation Lab is cheapest at $1.15 per 1,000. For the largest install base, Trudax Reddit Scraper Lite has 40,034 users at $4 per 1,000. Pushshift remains closed to the public, and the official Data API needs manual approval. ### What happened to the Reddit API in 2025 and 2026? Reddit closed the free paths in two steps. On November 11, 2025 it published the Responsible Builder Policy on r/redditdev and closed self-service access to the public Data API the same day; every new OAuth client now files a manual application covering use case, subreddits, and request volume (FetchLayer, 2026; Redditapis, 2026). Developers report multi-week waits and generic rejections. The second step landed in the last week of May 2026, when unauthenticated requests to any Reddit URL ending in .json began returning 403 Forbidden. FetchLayer dates the change to May 30; Crawlora reports the deprecation notice on May 28 and paraphrases Reddit's reasoning as stopping scraping without accountability and curbing agentic abuse. Either way, the last keyless way to read Reddit programmatically is gone. The historical archive went earlier. Pushshift lost public access on May 2, 2023 and has been restricted to Reddit moderators since (Xpoz, 2026). Reddit's 2023 commercial terms price approved access at about $0.24 per 1,000 calls, sold in blocks of roughly 50 million calls for about $12,000, which Crawlora and FetchLayer describe as a yearly minimum and Prowlo as monthly. - November 11, 2025: self-service OAuth app registration closed; every new Data API client needs manual approval (FetchLayer, 2026) - May 28 to 30, 2026: anonymous .json endpoints switched from JSON to 403 Forbidden (Crawlora, 2026; FetchLayer, 2026) - Since May 2, 2023: Pushshift restricted to Reddit moderators, no public API (Xpoz, 2026) - Official Data API: 100 queries per minute free for approved non-commercial clients; about $0.24 per 1,000 calls commercially, sold in blocks of about 50 million calls (Prowlo, 2026) ### Which Reddit API alternative is cheapest per 1,000 results? The table below uses free-plan prices from the Apify Store API on August 29, 2026. Datapika charges $2 per 1,000 posts, $3 per 1,000 posts with full comment trees, $2 per 1,000 search results, and $3 per 1,000 single post fetches, plus $0.00005 per run. It holds a 4.9 rating from 8 reviews with 188 users, and 0 of its 1,434 runs in the last 30 days failed. Trudax Reddit Scraper Lite leads on volume with 40,034 users, 7,088 in the last 30 days, a 4.57 rating from 39 reviews, and $4 per 1,000 results plus $0.02 per run ($3.40 on Gold); the apify/reddit-scraper URL redirects to it. Harshmaur costs $2 per 1,000 plus $0.02 per run, with 11,747 users and 4.48 from 39 reviews. Automation Lab is cheapest for posts alone at $1.15 per 1,000 posts, $0.575 per 1,000 comments and $0.003 per run, with 3,624 users and 4.8 from 6 reviews. Pratikdani is the outlier at $25 per 1,000 with 432 users and no reviews. Trudax, Harshmaur, and Automation Lab bill every comment as its own result; Datapika bills a post with comments as one $0.003 event whatever the thread length. - Cheapest bare post listings: Automation Lab at $1.15 per 1,000 posts (Apify Store, August 29, 2026) - Cheapest threads with comments: Datapika at $3 per 1,000 posts, comment count not billed - Most users: Trudax Reddit Scraper Lite, 40,034 total and 7,088 in the last 30 days - Most expensive per row: Pratikdani Reddit Post Scraper at $25 per 1,000 on a free plan, $15 on Gold - Only rows needing approval or login: the official Reddit Data API and Pushshift ### What does 10,000 Reddit posts actually cost? Take two real workloads. Workload A is 1,000 and 10,000 post listings from a subreddit feed with no comments. Workload B is 1,000 posts where each thread carries an average of 50 comments, typical for the hot tab of a mid-size subreddit. Workload A, 1,000 posts: Automation Lab $1.15 including its $0.003 start fee, Datapika $2.00, Harshmaur $2.02 including its $0.02 start fee, Trudax Lite $4.02, Pratikdani $25.00. At 10,000 posts: Automation Lab $11.50, Datapika $20.00, Harshmaur $20.02, Trudax Lite $40.02, Pratikdani $250.00. Automation Lab wins outright; Datapika edges Harshmaur for second only by the start fee. Workload B is 51,000 stored items on a per-result actor. Trudax Lite: 51,000 x $0.004 + $0.02 = $204.02. Harshmaur: 51,000 x $0.002 + $0.02 = $102.02. Automation Lab: 1,000 x $0.00115 + 50,000 x $0.000575 + $0.003 = $29.90. Datapika: 1,000 x $0.003 = $3.00. The official API would need roughly 1,000 comment-tree calls at about $0.24 per 1,000, so $0.24 metered, but only after approval and a block purchase guides put near $12,000. Residential proxy bandwidth, billed by Apify at $8 per GB, adds a few cents per 1,000 rows to every actor and is left out above. - 1,000 posts, no comments: Automation Lab $1.15, Datapika $2.00, Harshmaur $2.02, Trudax Lite $4.02, Pratikdani $25.00 - 10,000 posts, no comments: Automation Lab $11.50, Datapika $20.00, Harshmaur $20.02, Trudax Lite $40.02, Pratikdani $250.00 - 1,000 posts with 50 comments each: Datapika $3.00, Automation Lab $29.90, Harshmaur $102.02, Trudax Lite $204.02 - Official Data API for the same 1,000 threads: about $0.24 metered, plus approval and a reported $12,000 block minimum - All figures use free-plan Apify prices as of August 29, 2026 and exclude proxy bandwidth ### What data does each Reddit scraper return? Every Apify actor here returns the core post record: ID, permalink, subreddit, author, title, body, media URLs, score, comment count, and timestamp. Differences show up in scope and extras. Datapika covers six actions in one input schema: subreddit feeds sorted by hot, new, top, rising, or controversial with time filters; keyword search for posts or comments across Reddit or inside one subreddit; subreddit discovery by topic with mention counts; single post fetch with the full comment tree; and Reddit AI Answers, which returns a markdown answer with follow-up questions and source post IDs and appears in no other row. Trudax Lite adds user profiles and community metadata and accepts Reddit search URLs as input. Harshmaur bundles optional AI analysis at $0.0005 per analyzed item and custom labels at $0.0001 each. Automation Lab splits posts and comments into separate events, so you pull comments only when needed. Pratikdani returns posts only. The official Data API returns the richest object, including account-scoped vote and moderation fields, but only for approved clients at 100 queries per minute on the free tier. No live row reaches Pushshift's full-text history back to 2005; Datapika's archive fallback recovers some deleted content on a best-effort basis only. - Shared post fields: post_id, permalink, subreddit, author, title, body, media, score, num_comments, timestamp - Datapika only: Reddit AI Answers as markdown with follow_ups and source lists, $10 per 1,000 queries - Datapika comment trees: parent_id on every comment so nested threads can be rebuilt - Harshmaur only: in-run AI analysis at $0.0005 per item and custom labels at $0.0001 - Trudax Lite only: user profile and community metadata rows ### Which Reddit API alternative should you choose? Choose Datapika when you need whole threads, AI-agent access, or several Reddit tasks in one tool. The $3 per 1,000 posts-with-comments rate is the lowest in the table by roughly 10x for 50-comment threads, the actor is on the Apify MCP server so Claude or GPT agents can call it as a tool, and Reddit AI Answers is only available here. Its 188 users are the smallest install base of the paid rows, so weigh that against the price. Choose Automation Lab when you want the cheapest bare post listings and can accept a 6-review track record. Choose Harshmaur when in-run sentiment labels matter more than a slightly higher per-comment bill. Choose Trudax Reddit Scraper Lite when a 40,034-user history and a 4.57 rating outweigh paying double per result. Skip Pratikdani at $25 per 1,000; it is the most expensive row. Apply for the official Data API only if you need account-scoped fields or a contractual SLA and can wait through a manual review with no guaranteed outcome. Skip Pushshift unless you moderate a subreddit; for history, use community dumps such as Arctic Shift offline and a live scraper for anything newer. - Whole threads, MCP agents, or Reddit AI Answers: Datapika, $3 per 1,000 threads - Cheapest post-only feeds: Automation Lab, $1.15 per 1,000 posts - Sentiment tagging inside the scrape: Harshmaur, $2 per 1,000 plus $0.0005 per analyzed item - Largest user base and longest track record: Trudax Lite, 40,034 users since before the shutdown - Contract and SLA requirements: official Reddit Data API, if your application is approved ### Steps 1. Open the Datapika Reddit actor at https://apify.com/openclawai/reddit-scraper, or add it to an agent with https://mcp.apify.com/?tools=fetch-actor-details,openclawai/reddit-scraper. No Reddit account, OAuth app, or API key is needed. 2. Map your old API calls to actions: a subreddit listing becomes scrape_subreddit, a search query becomes search_posts or search_comments, a comments endpoint becomes fetch_post with a permalink, and switch includeComments on when you want full threads at $3 per 1,000. 3. Run with limit set to 50 to confirm the fields you need, then raise it toward the 500 cap; keep the default residential proxy group for anything beyond a test. 4. Export the dataset as JSON or CSV, or read it through the Apify dataset API, and schedule the saved input to replace the polling loop that used to hit the .json endpoints. ### FAQ Q: Is the Reddit API completely shut down in 2026? A: Not completely, but the self-serve path is closed. Since November 11, 2025 every new OAuth client must apply under the Responsible Builder Policy and wait for manual approval, and since the last week of May 2026 anonymous .json requests return 403 Forbidden (FetchLayer, 2026; Crawlora, 2026). Existing approved apps keep working at 100 queries per minute on the free non-commercial tier. For anyone without an approved app, a public-page scraper is the only working route. Q: What is the best Pushshift alternative in 2026? A: There is no live replacement for Pushshift's full-text history. Pushshift itself has been restricted to Reddit moderators since May 2, 2023 (Xpoz, 2026), and community archives such as Arctic Shift are offline dumps with search limited to one subreddit or user at a time; sources disagree on how current they are. For anything recent, use a live scraper: Datapika's search_comments action finds public comments by phrase at $2 per 1,000 results, and its archive fallback recovers some deleted posts on a best-effort basis. Q: How much does the official Reddit Data API cost compared with a scraper? A: Reddit's 2023 terms price commercial access at about $0.24 per 1,000 calls, cheaper per call than any scraper, but it is sold in blocks of about 50 million calls for roughly $12,000 and guides report rare approvals (Crawlora, 2026; Prowlo, 2026). A scraper bills per row with no contract: 10,000 posts cost $20.00 on Datapika, $20.02 on Harshmaur, and $40.02 on Trudax Lite as of August 29, 2026. Below a few million calls a month, the scraper is cheaper in practice. Q: Is it legal to scrape Reddit without the API? A: Scrapers in this comparison read only pages Reddit serves to anonymous visitors and do not bypass logins or access private subreddits. That keeps them outside the computer-access questions raised by authenticated scraping, but it does not remove your obligations under Reddit's user agreement, copyright in user posts, and privacy laws such as GDPR when you store usernames. Check the rules for your jurisdiction and purpose, and avoid republishing personal data at scale. Q: Why does Datapika charge $3 per 1,000 threads when others charge $2 per 1,000 results? A: Because the unit is different. Trudax, Harshmaur, and Automation Lab count every comment as a result, so a 1,000-thread run with 50 comments each stores 51,000 items and costs $204.02, $102.02, or $29.90 respectively. Datapika bills the post with its complete comment tree as one $0.003 event, so the same run costs $3.00 (Apify Store prices, August 29, 2026). For post-only feeds Datapika is $2 per 1,000, the same as Harshmaur and above Automation Lab. Q: Can an AI agent use a Reddit API alternative through MCP? A: Yes. Datapika's Reddit actor is exposed on the Apify MCP server at https://mcp.apify.com/?tools=fetch-actor-details,openclawai/reddit-scraper, so a Claude, GPT, or custom agent can read the input schema and price, then call scrape_subreddit, search_posts, or reddit_answers as tools and read the dataset back. Billing stays per result, so an agent pulling 300 search results spends $0.60. Trudax and Harshmaur are also callable through Apify's generic MCP tools, but neither offers the Reddit AI Answers action. ## Best Reddit scrapers that survived the 2026 API shutdown Canonical: https://datapika.com/compare/best-reddit-scrapers Updated: 2026-08-29 The best Reddit scraper in August 2026 depends on the job. For AI agents and full comment trees, Datapika's Reddit Scraper at $2 per 1,000 posts, $3 with comments, and a 4.9 rating is the strongest value, and the only actor here returning Reddit AI Answers. For the longest track record, Trudax's Reddit Scraper Lite has 40,036 users at $3.40 per 1,000. For bare posts at the lowest list price, Harsh Maur's Reddit Scraper lists from $1.50 per 1,000. All seven run without a Reddit API key, decisive once anonymous .json requests began returning 403 in late May 2026. ### What happened to the Reddit API in 2026, and why re-rank the scrapers now? Reddit closed its free data paths in two steps. On November 11, 2025, Reddit announced that self-service access to the public Data API was closed, so new OAuth apps now need manual approval that is rarely granted (FetchLayer, 2026). Crawlora dates Reddit's notice deprecating unauthenticated .json access to May 28, 2026, and FetchLayer reports that anonymous .json requests have returned 403 Forbidden since May 30, 2026 (Crawlora, 2026; FetchLayer, 2026). The commercial Data API still exists at roughly $0.24 per 1,000 requests, with enterprise agreements starting near $12,000 per year, and each request needs an approved app (Crawlora, 2026). Public HTML pages are the only route left that does not need Reddit's permission. That matters for any best-of list, because most rankings in Google's index were written before the cut. The most-cited one, dated May 26, 2026, names apify/reddit-scraper as the default pick, yet that URL now serves Trudax's Reddit Scraper Lite listing. Every number below was pulled from each actor's live Apify Store page on August 29, 2026, and four actors now list below the old $3.40 per 1,000 anchor. - November 11, 2025: self-service OAuth registration for Reddit's Data API closed (FetchLayer, 2026) - Late May 2026: anonymous .json requests return 403; Crawlora dates the notice to May 28, FetchLayer the cutoff to May 30 - Official commercial API: about $0.24 per 1,000 requests, enterprise from about $12,000 per year, approved app required (Crawlora, 2026) - All seven actors here read public Reddit pages and need no Reddit API key, OAuth token, or login (store pages, August 29, 2026) ### Which Reddit scraper is best in 2026? The ranked table, row by row Trudax's Reddit Scraper Lite is the incumbent by volume: 40,036 total users, 7,090 monthly, a 4.57 rating, and 94.6 percent run success, from $3.40 per 1,000 results. It covers posts, comments, communities, users, and media in one input, and its README still frames Reddit's API changes around April 2023. Datapika's Reddit Scraper is the smallest listing by users, 188, but has the highest review-backed rating, 4.9 from 8 reviews, and 100 percent success across 6,931 runs. It bills $2 per 1,000 posts, $3 per 1,000 with full comment trees, and $0.01 per Reddit AI Answer, the only AI Answers action in this comparison. Harsh Maur's is the strongest challenger on price and scale: from $1.50 per 1,000, 11,662 users with 3,346 monthly, a 4.48 rating, 98.9 percent success, and nested comments. labrat011's lists from $1.50 with a 5.00 rating, 98.4 percent success, and 296 users; its README says it parses server-rendered HTML because the .json API has returned 403 since May 2026. ParseForge's Reddit Posts Scraper: $3.00 per 1,000, 1,506 users, a 5.00 rating, 93.6 percent success. Trudax's full Reddit Scraper is a $45 per month rental rated 2.57; Pratik Dani's lists at $15.00 per 1,000, unrated. - Most users: Trudax Reddit Scraper Lite, 40,036 total and 7,090 monthly, 4.57 rating, from $3.40 per 1,000 - Highest review-backed rating and success rate: Datapika, 4.9 from 8 reviews, 100 percent of 6,931 runs - Lowest list price for posts: Harsh Maur and labrat011 from $1.50 per 1,000; Datapika is $2, Trudax Lite $3.40 - Only actor in this comparison returning Reddit AI Answers: Datapika, $0.01 per answer (August 29, 2026) ### How much does it cost to scrape 1,000 or 10,000 Reddit posts? Per-result pricing keeps the arithmetic simple, at store prices on August 29, 2026. With Datapika, 1,000 posts without comments cost $2.00 and 10,000 cost $20.00; with includeComments on, the same volumes cost $3.00 and $30.00 with every nested reply included. Trudax Reddit Scraper Lite at $3.40 per 1,000 charges $3.40 and $34.00. Harsh Maur's headline is from $1.50 per 1,000, but that is the discounted rate on higher Apify plans. Its README lists the base events as $0.02 per run plus $0.002 per item, so on the Free plan 1,000 posts cost $2.02 and 10,000 cost $20.02, level with Datapika. Trudax Lite and ParseForge discount by plan too. The rental model only wins at sustained volume. Trudax's $45 per month scraper breaks even against Lite at about 13,200 results per month ($45 divided by $3.40 per 1,000) and against Datapika's $2 rate at 22,500 posts, plus the compute the rental still bills. A daily brand sweep on Datapika of 500 search results ($1.00), 50 threads with full comments ($0.15), and 5 Reddit AI Answers ($0.05) costs $1.20 per day, about $36 per month. Residential proxy bandwidth is billed by Apify at $8 per GB on top of every actor here. - Datapika: 1,000 posts $2.00, 10,000 posts $20.00; with full comment trees $3.00 and $30.00 - Trudax Reddit Scraper Lite: $3.40 and $34.00; Harsh Maur on the Free plan: $2.02 and $20.02 ($0.02 per run plus $0.002 per item) - Trudax $45 per month rental breaks even at about 13,200 results per month versus Lite, 22,500 versus Datapika - Daily brand sweep on Datapika (500 search results, 50 full threads, 5 AI Answers): about $1.20 per day plus $8 per GB proxy ### Which Reddit scrapers return comments, search, subreddit discovery, and AI Answers? Post fields are close to interchangeable across the rated actors: title, body, author, score, comment count, timestamp, permalink, and media URLs appear in every one. The differences are in the actions each actor exposes and how deep the comment fetch goes. Datapika bundles six actions behind one input: scrape_subreddit, search_posts, search_comments, search_subreddits, fetch_post, and reddit_answers, with up to 20 parallel comment threads and a fallback chain through old.reddit.com and an archive that recovers some deleted content. Trudax Lite is the broadest generalist, with searchPosts, searchComments, searchCommunities, searchUsers, and searchMedia flags plus direct URLs. Harsh Maur covers subreddit feeds, keyword search with a searchComments option, post URLs, and nested comments over MCP. ParseForge is post-first, with roughly 30 fields per post including upvote ratio, flair, and virality scores, and top comments only as an add-on capped by commentLimit. labrat011 advertises full recursive comment trees with depth tracking. Comment search is not a Datapika exclusive: Trudax Lite and Harsh Maur both return matching comments by keyword, and Trudax Lite also discovers communities by search term. The one capability unique in this table is Datapika's reddit_answers action, which returns the markdown answer, follow-up questions, and source post IDs as a dataset row. - Full nested comment trees: Datapika, Trudax Lite, Harsh Maur, labrat011; ParseForge returns top comments only - Keyword comment search: Datapika, Trudax Lite, Harsh Maur; subreddit discovery: Datapika and Trudax Lite - Reddit AI Answers as a dataset row with sources and follow-ups: Datapika only, $0.01 per answer - MCP server access documented: Datapika, Trudax Lite, Trudax full, Harsh Maur, labrat011, ParseForge, Pratik Dani ### Which Reddit scraper should you choose? Choose Datapika when an agent or scheduled pipeline needs several Reddit actions behind one tool, when full comment trees matter, or when you want Reddit AI Answers alongside raw posts. At $2 and $3 per 1,000 it undercuts Trudax Lite by 41 percent on posts and 12 percent on comment-inclusive posts, and its 100 percent success rate across 6,931 runs is the highest in the table. Accept that 188 users is a fraction of the incumbents' base and that its rating rests on 8 reviews. Choose Trudax Reddit Scraper Lite when you want the most public usage history and the broadest input, covering communities, users, and media as well as posts and comments. You pay $3.40 per 1,000 and accept 94.6 percent success, the lowest among rated options in August 2026. Choose Harsh Maur's Reddit Scraper for bulk posts on a paid Apify plan, where its discounted rate from $1.50 per 1,000 beats everyone, if you do not need AI Answers; with 3,346 monthly users it is the best-adopted budget option. Consider labrat011 for the same list price with a post-shutdown build. Skip the $45 rental below roughly 13,000 results a month, and skip the $15 per 1,000 actor entirely. - Agents, comment trees, AI Answers, six actions in one tool: Datapika Reddit Scraper - Longest track record and broadest generalist input: Trudax Reddit Scraper Lite - Cheapest bulk posts on a paid Apify plan with meaningful adoption: Harsh Maur Reddit Scraper from $1.50 per 1,000 - Post-only pulls with about 30 fields: ParseForge at $3.00 per 1,000; rental only above about 13,000 results a month ### Steps 1. Open the Datapika Reddit Scraper at https://apify.com/openclawai/reddit-scraper with a free Apify account; no Reddit account or developer key is needed. Agents attach it through https://mcp.apify.com/?tools=fetch-actor-details,openclawai/reddit-scraper, where fetch-actor-details returns the schema and prices before any paid call. 2. Pick one action: scrape_subreddit, search_posts, search_comments, search_subreddits, fetch_post, or reddit_answers. Migrating from Trudax Lite or Harsh Maur, a subreddit-plus-sort input maps to scrape_subreddit and a keyword input to search_posts; leave includeComments off on the first run to stay at $2 per 1,000. 3. Run with limit 10 to 50, confirm the columns you need (post_id, permalink, subreddit_name, title, body, num_upvotes, num_comments, post_timestamp), then raise the limit toward 500 and enable includeComments where you need threads at $3 per 1,000. 4. Export JSON, CSV, or Excel, or read the dataset through the Apify API, then save the input as a schedule; 500 search results plus 50 full threads daily costs about $1.15 per day. ### FAQ Q: Can I still use the official Reddit API instead of a scraper in 2026? A: Only with Reddit's approval. Self-service registration for the Data API closed on November 11, 2025, and unauthenticated .json requests have returned 403 since May 30, 2026 (FetchLayer, 2026). Approved commercial access is quoted around $0.24 per 1,000 requests, and enterprise agreements start near $12,000 per year (Crawlora, 2026). Every actor in this comparison reads public Reddit pages instead, so it needs no key, but each is limited to what Reddit serves anonymously. Q: Which Reddit scrapers still work after the May 2026 .json shutdown? A: Any actor that parses public HTML pages rather than the .json endpoints. As of August 29, 2026, the Apify Store reports 100 percent run success for Datapika, 98.9 percent for Harsh Maur, 98.4 percent for labrat011, 94.6 percent for Trudax Lite, and 93.6 percent for ParseForge. Only labrat011's README names the May 2026 change explicitly; Datapika documents a fallback chain through old.reddit.com and an archive. Test anything below 95 percent before scheduling it. Q: Is it legal or allowed to scrape Reddit with these tools? A: These actors only read content Reddit serves to any logged-out visitor, and none bypasses authentication or reaches private subreddits. That does not remove your own obligations: Reddit's user agreement, copyright in user posts, and data-protection laws such as GDPR and CCPA govern how you store and use the rows, especially author names. Reddit's stated reason for the May 2026 change was scraping without accountability, so keep volumes proportionate, honor deletion requests, and review the rules in your jurisdiction before scaling. Q: Do any of these Reddit scrapers work with MCP and AI agents? A: All seven document Apify MCP server access, including Datapika, Trudax Lite, Harsh Maur, labrat011, ParseForge, and Pratik Dani. Datapika's single-tool URL is https://mcp.apify.com/?tools=fetch-actor-details,openclawai/reddit-scraper, and the paired fetch-actor-details tool lets an agent read the schema and per-result prices before running. Because billing is per row, an agent pulling 200 search results spends $0.40, and one reddit_answers call returning a synthesized answer with sources costs $0.01. Q: Why does Datapika rank highly with only 188 users? A: Because this ranking weighs price, success rate, and capability, not adoption alone. As of August 29, 2026 Datapika has 188 users against 40,036 for Trudax Lite, but a 4.9 rating from 8 reviews, 100 percent success across 6,931 runs, the lowest published rate for comment-inclusive posts at $3 per 1,000, and the only Reddit AI Answers action in this comparison. If a large user base is the reassurance you need, pick Trudax Lite or Harsh Maur; otherwise a 50-row Datapika test costs about 10 cents.