TheCrawler — Apify actor wrapper around the standalone `thecrawler` npm package. All scraping logic lives in the standalone engine; this is a thin
Thecrawler MCP server exists for a simple reason — assistants are far more useful when they can act on Thecrawler directly instead of describing what you should do. TheCrawler — Apify actor wrapper around the standalone thecrawler npm package. All scraping logic lives in the standalone engine; this is a thin charging + dataset-push shell. Version mirrors the engine version (0.3.3) it imports.
Scrape web pages, run LLM-powered structured extraction, or diagnose whether URLs are ready for a built-in extraction contract before spending LLM tokens. Open source engine (AGPL-3.0). $0.005 per successfully scraped page on Apify.
You will need 2 environment variables: THECRAWLER_LLM_API_KEY, THECRAWLER_API_KEY. The server will not start without them, which is usually why the tools fail to appear on a first run. Keep credentials in your client's env block or a secrets manager rather than in a file you might commit.
Being a remote server, there is no local install. You register the endpoint with your client, authorise it once, and the tools appear.
Among the search and retrieval options, the useful question is rarely "what can it do" but "what does it cost you to run" — permissions, credentials, and how much of your context its toolset consumes. It is maintained by manchittlab; worth a glance at recent repository activity before you build anything load-bearing on it.
SyncDev reviews every entry in this directory against the project's own documentation before publishing, and revisits them as servers change.
{
"mcpServers": {
"thecrawler": {
"command": "npx",
"args": ["-y", "npm"],
"env": {
"THECRAWLER_LLM_API_KEY": "your-value",
"THECRAWLER_API_KEY": "your-value"
}
}
}
}Add to claude_desktop_config.json, then restart Claude Desktop.
| Variable | Description | Required |
|---|---|---|
| THECRAWLER_LLM_API_KEY | Credential the server authenticates with. | Yes |
| THECRAWLER_API_KEY | Credential the server authenticates with. | Yes |
The simplest web tool that matters — fetch any URL and get model-ready markdown back.
Industrial-strength web extraction — render, scrape, crawl and search entire sites into clean markdown.
Puppeteer-powered browser control that drives pages from the accessibility tree instead of pixels.
Search built for AI, not humans — semantic web search that returns model-ready content, plus code context.
Answers, not links — delegate questions to Perplexity's search-grounded models and get cited responses back.
Put 6,000+ pre-built scrapers at your assistant's fingertips through one MCP endpoint.