Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:CrewAIInc CrewAI Scrapfly Tool

From Leeroopedia
Knowledge Sources
Domains Tools, Web_Scraping
Last Updated 2026-02-11 00:00 GMT

Overview

ScrapflyScrapeWebsiteTool scrapes web pages using the Scrapfly API with support for anti-bot bypass, headless browser rendering, and multiple output formats.

Description

ScrapflyScrapeWebsiteTool extends BaseTool and provides enterprise-grade web scraping through Scrapfly's infrastructure. ScrapflyScrapeWebsiteToolSchema defines the input schema with fields for url, scrape_format (literal "raw", "markdown", or "text", defaulting to "markdown"), optional scrape_config dictionary for advanced Scrapfly options, and ignore_scrape_failures flag. On initialization, it imports ScrapflyClient from the scrapfly package (with interactive installation via click.confirm if missing) and creates a client using the provided API key or SCRAPFLY_API_KEY environment variable. The _run() method constructs a ScrapeConfig object with the URL, format, and any additional config parameters, calls self.scrapfly.scrape(), and returns the content from the response. If ignore_scrape_failures is set, errors are logged and None is returned instead of raising.

Usage

Use this tool when agents need to scrape websites that are protected by anti-bot measures, require JavaScript rendering, or need proxy rotation. It is the enterprise-grade alternative to the simpler BeautifulSoup-based ScrapeWebsiteTool.

Code Reference

Source Location

  • Repository: CrewAI
  • File: lib/crewai-tools/src/crewai_tools/tools/scrapfly_scrape_website_tool/scrapfly_scrape_website_tool.py
  • Lines: 1-85

Signature

class ScrapflyScrapeWebsiteToolSchema(BaseModel):
    url: str = Field(description="Webpage URL")
    scrape_format: Literal["raw", "markdown", "text"] | None = Field(default="markdown")
    scrape_config: dict[str, Any] | None = Field(default=None)
    ignore_scrape_failures: bool | None = Field(default=None)

class ScrapflyScrapeWebsiteTool(BaseTool):
    name: str = "Scrapfly web scraping API tool"
    description: str = "Scrape a webpage url using Scrapfly and return its content as markdown or text"
    args_schema: type[BaseModel] = ScrapflyScrapeWebsiteToolSchema
    api_key: str | None = None
    scrapfly: Any | None = None
    package_dependencies: list[str]  # ["scrapfly-sdk"]
    env_vars: list[EnvVar]  # SCRAPFLY_API_KEY required

    def __init__(self, api_key: str)
    def _run(self, url: str, scrape_format="markdown",
             scrape_config=None, ignore_scrape_failures=None)

Import

from crewai_tools import ScrapflyScrapeWebsiteTool

I/O Contract

Inputs

Name Type Required Description
url str Yes URL of the webpage to scrape
scrape_format str or None No Output format: "raw", "markdown", or "text". Default "markdown"
scrape_config dict or None No Additional Scrapfly ScrapeConfig parameters
ignore_scrape_failures bool or None No If true, log errors and return None instead of raising

Outputs

Name Type Description
_run() returns str or None Scraped page content in the requested format, or None if failures are ignored

Usage Examples

Basic Usage

from crewai_tools import ScrapflyScrapeWebsiteTool

tool = ScrapflyScrapeWebsiteTool(api_key="your_scrapfly_api_key")
result = tool._run(url="https://example.com", scrape_format="markdown")

# With advanced config
result = tool._run(
    url="https://example.com",
    scrape_format="text",
    scrape_config={"render_js": True, "country": "US"},
    ignore_scrape_failures=True
)

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment