crawl4ai - Crawl JavaScript-heavy sites and extract structured data
Crawls JavaScript-rendered pages and multiple URLs, generates markdown, and extracts structured data with CSS or LLM-based strategies.
Tags
Updated: 2026-10-01Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Crawl JavaScript-heavy pages
- Generate markdown content
- Extract structured data
- Process multiple URLs concurrently
- Discover URLs from sitemaps
- Apply CSS extraction schemas
- Run LLM-based extraction
Inputs
- Target URLs
- CSS or JSON schemas
- Extraction prompts
- Browser configuration
- Crawler configuration
- URL lists
- Sitemap or domain sources
Outputs
- Markdown content
- Raw HTML
- Discovered links
- Media metadata
- Structured extracted data
- Batch output files
Requirements
- Installed crawl4ai package
- Configured Playwright browser
- Python runtime for SDK usage
- LLM access for LLM-based extraction
