Reviews / Firecrawl
Firecrawl Review
A web data API that turns sites into clean Markdown and JSON for AI agents
Last reviewed: October 2026
TL;DR
Firecrawl is the most widely adopted open-source way to turn web pages into LLM-ready Markdown or structured JSON, and it is the closest self-hostable match for web-reading APIs like Jina Reader. The self-hosted version leaves out Fire-engine, the cloud's anti-bot and rendering layer, so the open-source core does not match the cloud on hard-to-scrape sites.
License
AGPL-3.0
Self-hosted
Yes
Cloud option
Firecrawl Cloud at firecrawl.dev (free tier available)
GitHub stars
190.2K★
Checked October 2026
Category
Developer Tools & AI Infrastructure
Pricing
Free self-hosted core; Cloud has a free tier of 1,000 credits per month, then paid plans from $16/month billed yearly
What is Firecrawl?
Firecrawl is a web data API for AI systems. It searches the web, scrapes pages, crawls whole sites, and returns clean Markdown, HTML, screenshots, or JSON that follows a schema you define. It handles JavaScript-rendered pages, parses PDFs and DOCX files, and ships SDKs for Python, Node.js, Go, Rust, Java, and Elixir, plus a CLI and an MCP server for clients such as Cursor, Claude, and Windsurf.
Installation & self-hosting
Self-hosting uses Docker Compose from the repository and needs Git, Docker with Compose v2, and a free port 3002. The stack runs the API, PostgreSQL, Redis, RabbitMQ, workers, and a Playwright service, and it comes up on http://localhost:3002. The official guide says the baseline has no durable storage volumes, TLS, high availability, or authentication, so it is meant for trusted networks until you harden it. The guide does not publish a minimum host size.
User experience
Firecrawl is an API and SDK product rather than an app, so the experience is a request and a response: send a URL or a search query and receive Markdown or JSON. The hosted dashboard manages API keys and usage. We did not run the self-hosted stack ourselves; this section reflects the official documentation.
Key features
- Scrape any URL into Markdown, HTML, screenshots, metadata, or schema-based JSON
- Search that returns full-page Markdown with results, so no separate scrape step
- Interact mode for clicking, scrolling, typing, and multi-step flows such as logins
- Site crawling plus PDF and DOCX parsing
- MCP server, CLI, and SDKs for six languages
What Firecrawl does well
- By far the largest community in this space, with about 190,000 GitHub stars and daily commits
- Output designed for LLMs, which suits RAG pipelines, research agents, and data enrichment
- A generous free cloud tier and an easy path from hosted to self-hosted
Where Firecrawl falls short
- Self-hosted Firecrawl lacks Fire-engine, so screenshots, page actions, and advanced anti-bot handling are not in the default stack
- Agent, Browser, and some specialized formats are Cloud-only or need extra services
- AGPL-3.0 means a modified version offered over a network must share its source
- LLM-based extraction in a self-hosted setup needs an OpenAI-compatible provider or Ollama
Pricing
Self-hosting is free apart from your infrastructure. Cloud plans are credit-based: Free gives 1,000 credits per month with no card, Hobby is $16/month billed yearly for 5,000 credits, Standard is $83/month for 100,000, Growth is $333/month for 500,000, and Scale is $599/month for 1,000,000. Enterprise is custom.
Firecrawl vs Jina AI
Jina AI and Firecrawl both turn web pages into clean text for LLMs, but they are different products at the edges. Jina is a broader search foundation with embeddings, rerankers, and a reader API, while Firecrawl focuses on scraping, crawling, search, and structured extraction, and lets you run the core yourself. If you only need to read single URLs, Jina's reader is lighter; if you need crawling and structured output, Firecrawl goes further.
See all Jina AI alternativesBest for
- Teams building RAG pipelines, research agents, or lead-enrichment tools that need clean web data
- Developers who want a hosted API first and the option to self-host later
- Projects that need crawling and schema-based extraction, not just single-page reading
Not ideal for
- Sites that need heavy anti-bot handling, unless you use Firecrawl Cloud
- Teams that need managed browser sessions for automation, where a browser infrastructure product fits better
- Organizations that cannot accept AGPL-3.0 obligations on modified deployments
Alternatives
Final verdict
Firecrawl is the default choice for open-source web data for AI, with a community and release pace that no competitor matches. The honest caveat is that the best scraping features live in the cloud, so plan on Firecrawl Cloud if you scrape protected sites, and use the self-hosted core for simpler pages and internal use.