Skip to content

Crawl4AI

Open-source crawler that turns websites into LLM-ready Markdown

minylist score, out of 5

#7of 12 overall#2in Developer Tools
Cost
FreemiumFree to run on your own server
Hosting
Self-hostableYour data can stay on your servers
License
Apache-2.0Open source
Replaces
FirecrawlThe only one we've reviewed so far
On this page(13 sections)

Crawl4AI review: our verdict

The default free option for turning web pages into Markdown for an LLM, with more crawl controls than most paid APIs. Self-hosting means you run the browsers, proxies and scaling; the new cloud removes that but its launch prices may change.

An open-source Python crawler (Apache-2.0) that turns websites into clean Markdown and JSON for RAG and AI agents. Free to self-host; the hosted cloud is pay as you go with $10 of free credit at signup.

Best for

Python developers and AI builders who want LLM-ready web content without a monthly subscription, and are comfortable running their own crawler or paying per request in the cloud.

Where it stops

The library and Docker server are free, but every hard part (headless browsers, proxies, CAPTCHAs, retries, scaling) is yours to run. The cloud handles that and bills per credit at $0.001 each: a light scrape is 0.2 credits, a page needing a real browser is ten times that, a search is 0.5 credits, and LLM extraction adds token charges. The vendor labels these launch prices, discounted for the first year and subject to change.

What we like

  • Free and open source under Apache-2.0, with about 85,000 GitHub stars
  • Clean Markdown with boilerplate filters and numbered citations, built for LLMs
  • Deep crawling, adaptive crawling, sessions, proxies and stealth mode in one library
  • Docker server with a REST API, MCP endpoint and monitoring dashboard
  • Cloud credits never expire, and signup includes $10 of credit until December 31, 2026

Watch out for

  • Self-hosted, you handle browsers, proxies, blocks and scaling yourself
  • Python only for the library; other languages go through the Docker REST API or the cloud
  • Cloud pricing is marked as launch pricing that may change
  • Pages that need a real browser cost ten times a light fetch on the cloud

Crawl4AI screenshots

Crawl4AI homepage: 'Crawl4AI now available on Cloud', badges for 84,639 stars and the Apache-2.0 license, and a scrape box for api.crawl4ai.com

What is Crawl4AI?

Crawl4AI is an open-source web crawler written in Python that turns websites into clean Markdown and structured data for large language models. You give it a URL and get back Markdown with headings, tables and code intact, menus and footers filtered out, and links turned into a citation list.

It's aimed at developers building RAG pipelines, AI agents and data pipelines who don't want a monthly scraping bill. The library is free under Apache-2.0, and with about 85,000 GitHub stars it's one of the most-starred crawlers on GitHub. On Reddit it's the tool people recommend most when someone complains about paid scrapers' expiring credits.

There are now three ways to run it: as a library inside your Python code, as a Docker server you host, or through Crawl4AI Cloud, a hosted API launched in 2026 that runs the browsers and proxies for you and adds web search.

What it doesn't give you, self-hosted, is anything managed. Proxies, CAPTCHA handling, retries, scaling and uptime are your job. The cloud covers those, but it's new and its prices are labelled launch pricing.

A note on shelf life: Crawl4AI releases often (v0.9.4 landed on September 23, 2026), and the cloud's pricing is discounted for its first year. Check the README and the pricing page before you plan around either.

How Crawl4AI works

In the library, you open an AsyncWebCrawler and call arun with a URL. It drives a real browser (Chromium, Firefox or WebKit), runs JavaScript, waits for content and returns the page as Markdown, HTML, links, media and metadata. arun_many crawls many URLs with a memory-aware dispatcher.

Content filters decide what stays: pruning removes boilerplate, BM25 keeps passages that match a query, and an LLM filter can do it with a model. For structured data you define a CSS or XPath schema, or hand a schema to any LLM provider for typed JSON.

For whole sites, deep crawl strategies (breadth-first, depth-first, best-first) follow links with crash recovery, and the adaptive crawler stops once it has gathered enough to answer your query. URL seeding from sitemaps and Common Crawl finds pages before crawling.

The Docker server wraps the same engine in a REST API, MCP endpoint and dashboard, secured with your own token. Crawl4AI Cloud runs it for you behind one key: the engine is chosen per domain (cache, plain HTTP or browser), and each call returns its cost in credits in a response header.

Try Crawl4AI

Crawl4AI pricing

You can run Crawl4AI yourself for free under its Apache-2.0 license, so the real cost is the server and the hours someone spends keeping it updated.

The open-source library and Docker server are free forever under Apache-2.0. Your costs are servers, proxies if you need them, and LLM tokens if you use LLM extraction.

Crawl4AI Cloud is pay as you go in credits worth $0.001 each. A scrape is 0.2 credits at the light level; the effort multiplier is ×0.5 for pages already in Crawl4AI's archive, ×1 for light and standard fetches and ×10 when a real browser renders the page. A search is 0.5 credits, an answer 1 credit, extraction 0.2 credits plus tokens (2.5 credits in and 15 out per 5,000 tokens), and bulk jobs 0.2 credits per URL. Checked October 3, 2026.

New accounts get 10,000 credits ($10) once at signup until December 31, 2026, then $5 to start, and a 7-day pass of 20 credits exists for quick tests. Top-up packs are $10 for 10,000 credits, $25 for 27,500 and $100 for 125,000. Bought and signup credit never expire, and a paid top-up raises limits to 10 concurrent requests and 120 a minute.

In practice a light page costs about $0.0002 and a browser-rendered page about $0.002. Crawl4AI marks this as discounted launch pricing for the first year that may change, though credit you already hold keeps its value. If your volume is steady and you have someone to run servers, self-hosting is the cheapest option of all.

Prices change often. Check the current plans on the official site before you pay.

See Crawl4AI pricing

Crawl4AI features

  • Async Python crawler with Chromium, Firefox and WebKit
  • Markdown generation with pruning, BM25 and LLM content filters
  • Structured extraction with CSS, XPath, regex or any LLM into a JSON schema
  • Deep crawl (BFS, DFS, best-first) with crash recovery and adaptive crawling
  • Sessions, persistent browser profiles, proxies, stealth mode and hooks
  • Docker server with REST endpoints, MCP and a monitoring dashboard
  • Crawl4AI Cloud: scrape, search, answer, extract, batch and bulk jobs behind one API key

The depth is in crawl control: sessions with saved logins, hooks at every step, adaptive crawling that stops when it has enough, and Markdown filters you can tune per query. Most paid APIs hide these behind defaults. The trade is that you configure more, and on your own servers you also own the anti-bot fight.

How to set up Crawl4AI

  1. 1

    Install the library

    Install from PyPI and run the setup command once to download the browser.

    pip install -U crawl4ai
    crawl4ai-setup
  2. 2

    Crawl your first page

    Open an AsyncWebCrawler and print the Markdown for one URL to check the output quality on a site you care about.

    import asyncio
    from crawl4ai import AsyncWebCrawler
    
    async def main():
        async with AsyncWebCrawler() as crawler:
            result = await crawler.arun(url="https://news.ycombinator.com")
            print(result.markdown)
    
    asyncio.run(main())
  3. 3

    Tune the Markdown

    Add a pruning or BM25 content filter so only the main content reaches your model, and compare token counts before and after.

  4. 4

    Run it as a server

    Use the Docker image to get a REST API, an MCP endpoint and a dashboard, protected by your own API token, so other services and agents can call it.

  5. 5

    Add proxies for hard sites

    Configure proxy rotation and stealth mode for sites that block automation. Expect to pay a proxy provider for this when self-hosting.

  6. 6

    Or use the cloud instead

    Verify your email for a key and $10 of credit, then call api.crawl4ai.com/scrape. Use the estimate endpoint to see a call's credit range before running it at scale.

What people use Crawl4AI for

1

RAG over documentation

Deep-crawl a docs site, keep the fit Markdown and load it into a vector store. Use caching so re-crawls skip unchanged pages.

2

Self-hosted scraping for n8n

Run the Docker server and call it from n8n's HTTP Request node, as Reddit users do to avoid monthly credit plans. You still need proxies for protected sites.

3

Research agents

Give an agent the MCP endpoint so it can fetch and read pages. The adaptive crawler helps it stop once it has enough to answer.

4

Structured product data

Write a CSS schema once for a product page layout and extract thousands of pages with no LLM cost. Layout changes break schemas, so monitor failures.

5

Bursty or one-off crawls

Use the cloud's non-expiring credits for jobs that come in spikes. A browser-rendered page costs ten times a light fetch, so estimate first.

6

Crawling behind a login

Use a persistent browser profile with your saved session to crawl pages you have access to. Respect the site's terms; logged-in scraping carries more legal risk.

7

Training or evaluation datasets

Seed URLs from sitemaps or Common Crawl and crawl at scale on your own hardware. Check licensing of the content you collect before using it for training.

Crawl4AI vs Firecrawl

Most people weighing Crawl4AI are really asking whether to keep paying for Firecrawl. You trade convenience for control: run it yourself and the data is yours, or pay someone else to run it for you.

AspectCrawl4AIFirecrawl
Price to startFree (open source) · Cloud pay as you goFrom $19/month
Free planYesYes
Self-hostingYesNo
Your dataOn your own servers if you self-hostIn the vendor's cloud

Firecrawl figures are from its pricing page, checked 2026-10-02.

All Firecrawl alternatives

Crawl4AI next to similar tools

Crawl4AI vs Firecrawl
Firecrawl is a managed API from $19 a month for 5,000 credits that expire monthly, with search, interaction and a keyless MCP. Crawl4AI is free to self-host, and its cloud credits never expire. Firecrawl is less setup; Crawl4AI is more control and cheaper for bursty work.
Crawl4AI vs Apify
Apify's store has ready-made scrapers for specific sites, from $19 a month plus compute units. Crawl4AI is a general crawler you configure yourself. Use Apify when a maintained Actor exists; use Crawl4AI for general pages into an LLM.
Crawl4AI vs Scrapy
Scrapy is the long-standing Python crawling framework, fast for static HTML but without built-in browser rendering or LLM-ready Markdown. Crawl4AI is built for JavaScript sites and AI output.
Crawl4AI vs Playwright alone
Playwright gives you the browser and nothing else. Crawl4AI adds Markdown generation, content filtering, extraction strategies and crawl orchestration on top.

Bottom line

Our take on Crawl4AI

If you write Python and want web pages as Markdown for an LLM, install Crawl4AI and try it on your target sites today; it costs nothing. Move to the Docker server when other services need it, and try the cloud's free $10 if you'd rather not run browsers and proxies. Skip it if you want a fully managed API with search and support from day one; Firecrawl is the closer fit.

Get started with Crawl4AIFreemium · self-hostable · 4.1 / 5

Crawl4AI FAQ

12 questions
Is Crawl4AI free?

Yes. The library and the self-hosted Docker server are open source under Apache-2.0 and free forever, according to the README. Crawl4AI Cloud is paid per credit, with 10,000 free credits ($10) at signup until December 31, 2026.

How much does Crawl4AI Cloud cost?

One credit is $0.001. A light scrape is 0.2 credits, a page that needs a real browser is 2 credits (×10), a search is 0.5 credits and an answer 1 credit, with LLM tokens extra. Top-ups are $10 for 10,000 credits, $25 for 27,500 and $100 for 125,000. Checked October 3, 2026.

Do Crawl4AI Cloud credits expire?

No. The pricing page says signup credit and bought credit never expire; only the 7-day pass expires with its key. Results already in Crawl4AI's archive cost half.

Is Crawl4AI open source?

Yes, under the Apache-2.0 license, which allows commercial use and modification. The repository is github.com/unclecode/crawl4ai, and the latest release when we checked was v0.9.4 from September 23, 2026.

How do I install Crawl4AI?

Run pip install -U crawl4ai, then crawl4ai-setup once to install the browser. A Docker image (AMD64 and ARM64) runs it as a server with a REST API and MCP endpoint.

Does Crawl4AI work with Claude Code or Cursor?

Yes. The self-hosted Docker server exposes an MCP endpoint, and the cloud has one at api.crawl4ai.com/mcp that Claude Code, Codex, Cursor and OpenCode can add with an API key.

Can Crawl4AI scrape JavaScript-heavy sites?

Yes. It runs a real browser, can execute JavaScript, wait for elements and scroll full pages for infinite scroll. Stealth mode and an undetected-browser adapter help with sites that detect automation, but proxies are yours to supply when self-hosting.

Can Crawl4AI extract structured data?

Yes. Use CSS, XPath or regex schemas with no LLM, or LLMExtractionStrategy with any provider for typed JSON. The cloud's /extract endpoint does it with a plain-English instruction and no LLM key of your own.

Who makes Crawl4AI?

It was created by the developer known as UncleCode, whose GitHub profile names them as the author of Crawl4AI and founder of Kidocode. They say in the README that they built it in 2023 after a paid web-to-Markdown tool under-delivered.

Crawl4AI vs Firecrawl: which should I use?

Crawl4AI is free to self-host and its cloud credits never expire. Firecrawl is a managed API with monthly plans from $19 for 5,000 credits, plus search, interaction and a keyless MCP. Pick Crawl4AI for control and bursty volume, Firecrawl for the least setup.

Is Crawl4AI good for RAG?

Yes. Its Markdown keeps headings, tables and code, can filter boilerplate, and turns links into a numbered citation list, which helps retrieval and source attribution.

Does Crawl4AI handle proxies and CAPTCHAs?

Self-hosted, it supports proxies with authentication and rotation, but you bring the proxies and nothing solves CAPTCHAs for you. The cloud says it handles bot walls and proxies automatically.

How we review tools

We read the pricing page, the license and the self-hosting docs for every listing, then write down what it costs and where it stops. Vendors can't buy a better score or a higher spot. Some buttons are affiliate links, and the note at the top of the page says so. How affiliate links work here

More in Developer Tools

Browse Developer Tools

Ask AI about minylist

Open your assistant with a ready-made question about the site.