Reviews / Firecrawl

Firecrawl Review

A web data API that turns sites into clean Markdown and JSON for AI agents

Last reviewed: October 2026

TL;DR

Firecrawl is the most widely adopted open-source way to turn web pages into LLM-ready Markdown or structured JSON, and it is the closest self-hostable match for web-reading APIs like Jina Reader. The self-hosted version leaves out Fire-engine, the cloud's anti-bot and rendering layer, so the open-source core does not match the cloud on hard-to-scrape sites.

License

AGPL-3.0

Self-hosted

Yes

Cloud option

Firecrawl Cloud at firecrawl.dev (free tier available)

GitHub stars

190.2K★

Checked October 2026

Category

Developer Tools & AI Infrastructure

Pricing

Free self-hosted core; Cloud has a free tier of 1,000 credits per month, then paid plans from $16/month billed yearly

What is Firecrawl?

Firecrawl is a web data API for AI systems. It searches the web, scrapes pages, crawls whole sites, and returns clean Markdown, HTML, screenshots, or JSON that follows a schema you define. It handles JavaScript-rendered pages, parses PDFs and DOCX files, and ships SDKs for Python, Node.js, Go, Rust, Java, and Elixir, plus a CLI and an MCP server for clients such as Cursor, Claude, and Windsurf.

Installation & self-hosting

Self-hosting uses Docker Compose from the repository and needs Git, Docker with Compose v2, and a free port 3002. The stack runs the API, PostgreSQL, Redis, RabbitMQ, workers, and a Playwright service, and it comes up on http://localhost:3002. The official guide says the baseline has no durable storage volumes, TLS, high availability, or authentication, so it is meant for trusted networks until you harden it. The guide does not publish a minimum host size.

User experience

Firecrawl is an API and SDK product rather than an app, so the experience is a request and a response: send a URL or a search query and receive Markdown or JSON. The hosted dashboard manages API keys and usage. We did not run the self-hosted stack ourselves; this section reflects the official documentation.

Key features

  • Scrape any URL into Markdown, HTML, screenshots, metadata, or schema-based JSON
  • Search that returns full-page Markdown with results, so no separate scrape step
  • Interact mode for clicking, scrolling, typing, and multi-step flows such as logins
  • Site crawling plus PDF and DOCX parsing
  • MCP server, CLI, and SDKs for six languages

What Firecrawl does well

  • By far the largest community in this space, with about 190,000 GitHub stars and daily commits
  • Output designed for LLMs, which suits RAG pipelines, research agents, and data enrichment
  • A generous free cloud tier and an easy path from hosted to self-hosted

Where Firecrawl falls short

  • Self-hosted Firecrawl lacks Fire-engine, so screenshots, page actions, and advanced anti-bot handling are not in the default stack
  • Agent, Browser, and some specialized formats are Cloud-only or need extra services
  • AGPL-3.0 means a modified version offered over a network must share its source
  • LLM-based extraction in a self-hosted setup needs an OpenAI-compatible provider or Ollama

Pricing

Self-hosting is free apart from your infrastructure. Cloud plans are credit-based: Free gives 1,000 credits per month with no card, Hobby is $16/month billed yearly for 5,000 credits, Standard is $83/month for 100,000, Growth is $333/month for 500,000, and Scale is $599/month for 1,000,000. Enterprise is custom.

Firecrawl vs Jina AI

Jina AI and Firecrawl both turn web pages into clean text for LLMs, but they are different products at the edges. Jina is a broader search foundation with embeddings, rerankers, and a reader API, while Firecrawl focuses on scraping, crawling, search, and structured extraction, and lets you run the core yourself. If you only need to read single URLs, Jina's reader is lighter; if you need crawling and structured output, Firecrawl goes further.

See all Jina AI alternatives

Best for

  • Teams building RAG pipelines, research agents, or lead-enrichment tools that need clean web data
  • Developers who want a hosted API first and the option to self-host later
  • Projects that need crawling and schema-based extraction, not just single-page reading

Not ideal for

  • Sites that need heavy anti-bot handling, unless you use Firecrawl Cloud
  • Teams that need managed browser sessions for automation, where a browser infrastructure product fits better
  • Organizations that cannot accept AGPL-3.0 obligations on modified deployments

Alternatives

Jina AIBrowserbaseCrawl4AI
Compare them all

Final verdict

Firecrawl is the default choice for open-source web data for AI, with a community and release pace that no competitor matches. The honest caveat is that the best scraping features live in the cloud, so plan on Firecrawl Cloud if you scrape protected sites, and use the self-hosted core for simpler pages and internal use.

Sources