WebCrawler API

WebCrawler API

Hosted crawling and extraction API that turns any URL into clean markdown, HTML, or structured JSON for AI agents and RAG pipelines.

78/100Safe BetFree · from $29/moFreemium

If you're feeding web content into a RAG pipeline or an AI agent and don't want to run your own headless-browser fleet, WebCrawler API is one of the least annoying ways to get clean markdown out of messy pages. The 2026 releases matter more than the marketing: Wagent (June 19, 2026) turns a natural-language prompt into structured JSON with a required max_spend_usd cap; the /v2/scrape output_formats array (April 14, 2026) returns markdown, cleaned, html and links in one call; and free cache hits (September 2, 2026) cut repeat-crawl cost. Compare it to Firecrawl if you want a broader managed-scraping ecosystem, or to running Playwright plus Readability yourself if volume is low and you'd

Verified 6d ago · liveness 78/100 · cite: rightaichoice.com/tools/webcrawler-api

Best for
  • Developers building AI support bots or knowledge products from web content
  • Data scientists assembling RAG pipelines that need clean, chunkable markdown
  • Product and growth teams monitoring competitor pricing, docs and feature pages
  • No-code operators wiring web data into Zapier, Make or n8n
Not ideal for
  • Teams that need to self-host and audit the scraping code themselves
  • Projects that must stream page content in real time as it is fetched
  • Very high-volume crawl jobs that exceed the parallel-request ceiling of the plan you buy
Visit Website

IntermediateA developer with an API key from the dashboard can send a first /v2/scrape call and read markdown within about 15 minutes, since the homepage quickstart and docs cover cURL, Node.js, Python, PHP and Java. Wiring a full crawl job with URL filters plus webhooks is a half-day. No-code operators on Zapier, Make, n8n or Integrately take under an hour using the published platform guides. Self-serveWeb · API · CLIAPI availableVerified 6d ago
Pricing
Free · from $29/mo
FreemiumFree tier4 plans5 hidden costs
Learning curve
Intermediate
A developer with an API key from the dashboard can send a first /v2/scrape call and read markdown within about 15 minutes, since the homepage quickstart and docs cover cURL, Node.js, Python, PHP and Java. Wiring a full crawl job with URL filters plus webhooks is a half-day. No-code operators on Zapier, Make, n8n or Integrately take under an hour using the published platform guides. Self-serve
Runs on
WebAPICLI
API available · 6 integrations
Who it's for
Developer building an AI support botGrowth analyst monitoring competitor pricingData enrichment operator filling partial company records
Live sentiment
Is WebCrawler API actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip WebCrawler API if you need to self-host and audit the crawler's code, must stream page content in real time as it is fetched, or need enterprise SLAs, a named support contact and a formal security review.

The 30-second take
Biggest gripe

Wagent runs bill per LLM token plus per page scraped, so a complex task can run from a fraction of a cent up to the max_spend_usd cap you set — set it low on first runs.

Price reality

WebCrawler API's per-page rate falls as you commit: $0.002/page pay-as-you-go with no monthly commitment, $0.0015/page on the $99/mo Professional plan, and $0.001/page on the $499/mo Business plan, with LLM-cleaned markdown priced from $0.006/page down to $0.005/page. That suits solo builders and small teams doing hundreds to low thousands of pages a month; above that the $499/mo tier's 50 parallel requests and $0.001/page rate is the one that pays for itself. High-volume shops that already run

In short

WebCrawler API — Hosted crawling and extraction API that turns any URL into clean markdown, HTML, or structured JSON for AI agents and RAG pipelines. Best for Developers building AI support bots or knowledge products from web content, Data scientists assembling RAG pipelines that need clean, chunkable markdown, Product and growth teams monitoring competitor pricing, docs and feature pages. Free to start; paid plans from $29/mo.

What's new in WebCrawler API

Checked 6 days ago

Across the latest 5 updates: 1 feature update, 2 launches, 1 pricing change and 1 changelog entry.

Viability Score

78/100
Safe Bet

How well maintained and how widely used is WebCrawler API? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
not measured
Site health
95
User sentiment
not measured
What the vendor publishes
60

Last calculated: October 2026

How we score →

Key Features

  • Markdown extraction with menus, cookie banners, ads and footers stripped
  • LLM-cleaned /markdown endpoint (1–2s added latency, raw-markdown fallback)
  • /v2/scrape output_formats array: markdown, cleaned, html, links in one call
  • Structured 'links' output returning all page hyperlinks as an array
  • Structured extraction with JSON Schema output
  • Crawling Agent (Wagent) returns structured JSON from a natural-language prompt
  • Wagent model selection incl. openai/gpt-5.4-mini, anthropic/claude-sonnet-4.6, google/gemini-3.1-flash-lite-preview
  • Required max_spend_usd spending cap on every Wagent run
  • Change detection feeds with full content, additions, removals and diffs
  • Free cache hits on matching scrape requests and matching Wagent runs
  • Smart caching returning frequent pages in ~0.9s instead of ~4.7s (max_age=0 to bypass)
  • Sitemap-assisted crawl discovery via automatic sitemap.xml parsing
  • Synchronous /v2/scrape endpoint with a 3-minute timeout
  • Proxies, retries, headless browsers, JavaScript rendering, CAPTCHA solving, anti-bot bypass
  • Official SDKs for JavaScript, Python, PHP, Java and .NET

About WebCrawler API

FreemiumIntermediateAPI availableWeb · API · CLI

WebCrawler API is a hosted web crawling and data extraction service that returns page content in a form AI systems can consume directly. Send a URL or a sitemap to the scrape or crawl endpoints and you get back markdown with menus, cookie banners, footers and ads already stripped, so the output drops into a prompt or a vector store without a cleanup pass. The /v2/scrape endpoint accepts an output_formats array so you can request markdown, cleaned, html, and links in the same call — the links format returns hyperlinks as a structured array for link-graph or SEO work. Proxies, retries, headless browsers, JavaScript rendering, rate-limit handling, CAPTCHA solving and anti-bot bypass all run server-side, with each request routed through the fastest path that can actually fetch the page. Smart caching returns frequently requested pages in about 0.9 seconds instead of 4.7, and since September 2, 2026 cache hits are free, including Wagent runs that match a prior model, prompt and URL set. Change detection feeds deliver only changed pages with full content, additions, removals and diffs, so you stop re-crawling unchanged docs. Structured extraction uses JSON Schema, and the Crawling Agent (Wagent, launched June 19, 2026) takes a natural-language prompt plus seed URLs, browses and follows relevant links, and returns structured JSON — with a required max_spend_usd cap per run and billing per LLM token plus per page scraped. Wagent lets you choose a model; the changelog names openai/gpt-5.4-mini, anthropic/claude-sonnet-4.6 and google/gemini-3.1-flash-lite-preview among the options. The newer /markdown endpoint adds an LLM cleaning pass (1–2 seconds of extra latency, falling back to raw markdown if the cleaning call fails) and is priced at $0.006/page on Starter, $0.0055/page on Professional and $0.005/page on Business. Official SDKs cover JavaScript, Python, PHP, Java and .NET, and no-code paths run through Zapier, Make, n8n, Integrately and an MCP server. Pricing is per page from $0.002 on pay-as-you-go with no monthly commitment, dropping to $0.0015/page on the $99/mo Professional plan and $0.001/page on the $499/mo Business plan; a free Chrome extension (September 17, 2026) converts the current page to markdown client-side with no API key.

Behind the Verdict

WebCrawler API sits in the managed-scraping layer, and its differentiation is output quality plus the 2026 agent features rather than raw crawl throughput. The core loop is simple: POST a URL to /v2/scrape with output_formats: ["markdown", "links"] and get back both the cleaned article body and a structured hyperlink array. Cleaning is not a post-processing afterthought — menus, cookie banners, footers and ads are removed server-side, and the vendor's own FAQ claims the markdown is clean enough to pass straight into an LLM prompt or vector store. That's the claim to test on your own pages, but the pipeline design is right. The strongest additions of 2026 are Wagent and change detection. Wagent (June 19, 2026) is an agent that performs web searches, reads pages, decides which links are relevant, follows them and extracts data into a JSON Schema you supply — you give it a prompt, seed URLs, an output_schema, a model choice (the changelog lists openai/gpt-5.4-mini, anthropic/claude-sonnet-4.6 and google/gemini-3.1-flash-lite-preview) and a required max_spend_usd cap. Billing is per LLM token plus per page scraped, so the spending cap isn't optional garnish; it's the control that keeps a sprawling agent run from surprising you. Change detection feeds take the opposite approach to freshness: instead of a polling loop, you subscribe to a site and receive only changed pages with diffs, which is the correct shape for keeping a documentation knowledge base current without burning tokens. Operationally, the boring parts are handled: residential proxies, automatic retries, rate-limit handling, real headless browsers, JavaScript rendering, CAPTCHA solving and anti-bot bypass, with fastest-path selection per request. The published numbers are 98% extraction success, 2.5s average crawling time and 99.98% uptime, with cached pages around 0.9s versus 4.7s uncached — pass max_age=0 when you must bypass cache. Job cost is visible per job in the dashboard and filterable by status, and GET /organization/balance lets you check spendable balance programmatically. Where it strains. The /markdown LLM-cleaning endpoint adds one to two seconds of latency per page and silently falls back to raw uncleaned markdown if the LLM step fails — you need to detect that fallback if clean output is a hard requirement. The synchronous /v2/scrape endpoint times out after three minutes (raised from two in May 2026), so the heaviest pages still belong on the async job path or webhooks. Parallel request ceilings scale with plan — 5 on pay-as-you-go, 10 on the $29/mo Starter, 20 on the $99/mo Professional, 50 on the $499/mo Business — and that ceiling, not the per-page rate, is usually what limits a large backfill. This is also an indie-built product with support handled directly by the founder, which is fine for a developer tool and less fine if you need a named support contact or a formal security review. And the credit model means you're prepaying for a service whose upstream target

Researching WebCrawler API? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas WebCrawler API actually fits — and what changes day-one when you adopt it.

Developer building an AI support bot

Points the crawl endpoint at the product's docs domain, gets back cleaned markdown with nav and footers removed, chunks it and loads it into a vector store, then subscribes to a change detection feed on the same site.

Outcome: The knowledge base stays current from feed diffs instead of scheduled full re-crawls, and clean markdown means no pre-ingest cleanup script.

Growth analyst monitoring competitor pricing

Sets up change detection feeds on five competitor pricing pages, each delivery carrying full content plus the diff, and uses the 'links' output format to seed crawl lists for feature pages.

Outcome: Pricing moves surface the day the vendor page changes, with the new value and the old one in the same payload instead of a manual diff.

Data enrichment operator filling partial company records

Sends partial company records to Wagent with a natural-language prompt, seed URLs, an output_schema for the missing fields and a low max_spend_usd cap, so the agent searches, follows links and returns structured JSON.

Outcome: Missing fields come back sourced and schema-valid, and the spending cap keeps an ambiguous run bounded.

Use Cases

Models Under the Hood

openai/gpt-5.4-minianthropic/claude-sonnet-4.6google/gemini-3.1-flash-lite-preview

as of 2026-09-24

Limitations

  • The /markdown endpoint's LLM cleaning step adds one to two seconds of latency per page and falls back to raw, uncleaned markdown if the LLM call fails, so you must detect that fallback if clean output is mandatory.
  • The synchronous /v2/scrape endpoint times out after three minutes — heavy pages belong on the async job path or webhooks.
  • Parallel request ceilings scale with plan (5 pay-as-you-go, 10 Starter, 20 Professional, 50 Business), and that ceiling rather than the per-page rate usually caps a large backfill.
  • Change detection feeds deliver pages only as the service re-checks them, and cache-hit savings require repeat requests of the same URL.
  • It is an indie-built product with support handled directly by the founder.

as of 2026-10-02

Verification history

We have re-verified WebCrawler API 6 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published WebCrawler API tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Pay as you go

$0/mo

Ideal for

Solo builders and prototype teams who want to test crawl quality on real sites before committing to a monthly plan.

What this tier adds

Starting tier: no monthly commitment, $0.002/page from the first request, up to 5 parallel requests, unlimited proxy included.

Starter

$29/mo

Ideal for

Small projects with steady but modest crawling needs that want a fixed monthly allowance instead of per-request top-ups.

What this tier adds

Adds a monthly credit allowance and raises the parallel request ceiling from 5 to 10; LLM-cleaned markdown is $0.006/page.

Professional

$99/mo

Ideal for

Growing teams running continuous crawls alongside change detection feeds and Wagent runs.

What this tier adds

Per-page rate drops to $0.0015, LLM-cleaned markdown to $0.0055/page, parallel requests double to 20, and up to 20 API keys per organization.

Business

$499/mo

Ideal for

High-volume crawling operations where the per-page rate and the parallel request ceiling are the deciding factors.

What this tier adds

Lowest rates at $0.001/page and $0.005/page for LLM-cleaned markdown, with 50 parallel requests and the organization balance endpoint.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Wagent runs bill per LLM token plus per page scraped, so a complex task can run from a fraction of a cent up to the max_spend_usd cap you set — set it low on first runs.
  • LLM-cleaned markdown costs more than plain scrape output: $0.006/page on Starter, $0.0055/page on Professional and $0.005/page on Business, versus a $0.002–$0.001/page scrape. Use POST /scrape with
  • Parallel requests are capped by plan — 5 pay-as-you-go, 10 Starter, 20 Professional, 50 Business — so a large backfill can queue even when your credit balance is healthy.
  • Each page attempted in a crawl consumes credits; the 'only pay for successful requests' guarantee applies to successful requests, not to a job whose target site has changed shape.
  • Top-up credits are a separate purchase from your subscription allowance, so an included-allowance overrun means another checkout before the next batch of jobs finishes.

Where the pricing makes sense

The company stage and team size where WebCrawler API's pricing actually pencils out — and where peers do it cheaper.

WebCrawler API's per-page rate falls as you commit: $0.002/page pay-as-you-go with no monthly commitment, $0.0015/page on the $99/mo Professional plan, and $0.001/page on the $499/mo Business plan, with LLM-cleaned markdown priced from $0.006/page down to $0.005/page. That suits solo builders and small teams doing hundreds to low thousands of pages a month; above that the $499/mo tier's 50 parallel requests and $0.001/page rate is the one that pays for itself. High-volume shops that already run

Setup time & first value

How long it actually takes to get something useful out of WebCrawler API — broken out by persona, not the marketing-page minute.

A developer with an API key from the dashboard can send a first /v2/scrape call and read markdown within about 15 minutes, since the homepage quickstart and docs cover cURL, Node.js, Python, PHP and Java. Wiring a full crawl job with URL filters plus webhooks is a half-day. No-code operators on Zapier, Make, n8n or Integrately take under an hour using the published platform guides. Self-serve

Switching to or from WebCrawler API

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From a self-hosted Playwright/Puppeteer scraper: replace the browser fleet with /v2/scrape output_formats:["markdown"] and drop proxy, retry and CAPTCHA handling entirely.
  • →From Firecrawl or a similar managed scraper: keep your ingestion code and swap the base URL, since output is markdown plus structured JSON Schema output.
  • →From manual copy-paste content workflows: install the free Chrome extension to convert a single page to markdown while you script the rest of the pipeline.
  • →From a scheduled full re-crawl: move the same doc sites onto change detection feeds so only changed pages are delivered with diffs.
Migrating out
  • ↗To a self-hosted Playwright plus Readability stack: export crawl lists via the links output format and reimplement cleaning yourself for full code control.
  • ↗To an LLM vendor's built-in web tool: send seed URLs from the crawl output into the vendor's retrieval feature, accepting less control over cleaning quality.
  • ↗To a no-code-first platform: keep Zapier, Make or n8n as the orchestration layer and swap the crawl step behind the same automation.

Integrations

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “WebCrawler API”, and we withheld 6: 6 did not mention WebCrawler API. We are showing none, because we could not prove any of them are about WebCrawler API.

Tools that pair well with WebCrawler API

Common stack mates teams adopt alongside WebCrawler API, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to WebCrawler API

View all
Spider Cloud

Spider Cloud

Spider Cloud is a web scraping and crawling API that turns live pages into markdown or JSON for agents and RAG pipelines.

FreemiumTry
Context.dev

Context.dev

Web scraping API that turns any page into clean Markdown, structured JSON, screenshots, and brand data for AI agents

FreemiumTry
Crawl4AI

Crawl4AI

Open-source, LLM-ready crawler that turns any URL into clean Markdown, typed JSON, or search results — self-hosted or via a paid cloud API.

FreemiumTry

Frequently Asked Questions

Used WebCrawler API? Help shape our editorial sentiment research.