Implementation:CrewAIInc CrewAI Scrapfly Tool
| Knowledge Sources | |
|---|---|
| Domains | Tools, Web_Scraping |
| Last Updated | 2026-02-11 00:00 GMT |
Overview
ScrapflyScrapeWebsiteTool scrapes web pages using the Scrapfly API with support for anti-bot bypass, headless browser rendering, and multiple output formats.
Description
ScrapflyScrapeWebsiteTool extends BaseTool and provides enterprise-grade web scraping through Scrapfly's infrastructure. ScrapflyScrapeWebsiteToolSchema defines the input schema with fields for url, scrape_format (literal "raw", "markdown", or "text", defaulting to "markdown"), optional scrape_config dictionary for advanced Scrapfly options, and ignore_scrape_failures flag. On initialization, it imports ScrapflyClient from the scrapfly package (with interactive installation via click.confirm if missing) and creates a client using the provided API key or SCRAPFLY_API_KEY environment variable. The _run() method constructs a ScrapeConfig object with the URL, format, and any additional config parameters, calls self.scrapfly.scrape(), and returns the content from the response. If ignore_scrape_failures is set, errors are logged and None is returned instead of raising.
Usage
Use this tool when agents need to scrape websites that are protected by anti-bot measures, require JavaScript rendering, or need proxy rotation. It is the enterprise-grade alternative to the simpler BeautifulSoup-based ScrapeWebsiteTool.
Code Reference
Source Location
- Repository: CrewAI
- File: lib/crewai-tools/src/crewai_tools/tools/scrapfly_scrape_website_tool/scrapfly_scrape_website_tool.py
- Lines: 1-85
Signature
class ScrapflyScrapeWebsiteToolSchema(BaseModel):
url: str = Field(description="Webpage URL")
scrape_format: Literal["raw", "markdown", "text"] | None = Field(default="markdown")
scrape_config: dict[str, Any] | None = Field(default=None)
ignore_scrape_failures: bool | None = Field(default=None)
class ScrapflyScrapeWebsiteTool(BaseTool):
name: str = "Scrapfly web scraping API tool"
description: str = "Scrape a webpage url using Scrapfly and return its content as markdown or text"
args_schema: type[BaseModel] = ScrapflyScrapeWebsiteToolSchema
api_key: str | None = None
scrapfly: Any | None = None
package_dependencies: list[str] # ["scrapfly-sdk"]
env_vars: list[EnvVar] # SCRAPFLY_API_KEY required
def __init__(self, api_key: str)
def _run(self, url: str, scrape_format="markdown",
scrape_config=None, ignore_scrape_failures=None)
Import
from crewai_tools import ScrapflyScrapeWebsiteTool
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| url | str | Yes | URL of the webpage to scrape |
| scrape_format | str or None | No | Output format: "raw", "markdown", or "text". Default "markdown" |
| scrape_config | dict or None | No | Additional Scrapfly ScrapeConfig parameters |
| ignore_scrape_failures | bool or None | No | If true, log errors and return None instead of raising |
Outputs
| Name | Type | Description |
|---|---|---|
| _run() returns | str or None | Scraped page content in the requested format, or None if failures are ignored |
Usage Examples
Basic Usage
from crewai_tools import ScrapflyScrapeWebsiteTool
tool = ScrapflyScrapeWebsiteTool(api_key="your_scrapfly_api_key")
result = tool._run(url="https://example.com", scrape_format="markdown")
# With advanced config
result = tool._run(
url="https://example.com",
scrape_format="text",
scrape_config={"render_js": True, "country": "US"},
ignore_scrape_failures=True
)