What it is
Scrapling is an adaptive web scraping framework that covers everything from a single request to a full-scale crawl. Its parser relocates elements when pages change. Its fetchers can bypass anti-bot systems like Cloudflare Turnstile. Its Scrapy-like spider framework supports concurrent multi-session crawls with pause/resume, proxy rotation and AutoThrottle.
Who it's for
- Web scrapers who want selectors that survive website design changes
- Developers building concurrent, resumable crawls in Python
- Teams feeding AI agents or RAG pipelines with scraped pages, through the MCP server, Agent Skill or Markdown output
- Scrapy or BeautifulSoup users looking for a familiar API
Requirements
Requirements
- Python (the README links a supported-versions badge but states no version)
- A Playwright Chromium or Google Chrome browser for DynamicFetcher browser automation
Examples
Adaptive fetching and selection
pythonfrom scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher
StealthyFetcher.adaptive = True
p = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True) # Fetch website under the radar!
products = p.css('.product', auto_save=True) # Scrape data that survives website design changes!
products = p.css('.product', adaptive=True) # Later, if the website structure changes, pass `adaptive=True` to find them!What it does: Fetches a page with StealthyFetcher, saves the matched elements with auto_save=True, and relocates them later with adaptive=True if the site structure changes.
Scaling up to a full crawl with a Spider
pythonfrom scrapling.spiders import Spider, Response
class MySpider(Spider):
name = "demo"
start_urls = ["https://example.com/"]
async def parse(self, response: Response):
for item in response.css('.product'):
yield {"title": item.css('h2::text').get()}
MySpider().start()What it does: Defines a Scrapy-like spider with start_urls and an async parse callback that yields items, then starts the crawl.
Pros & cons
Pros
- Pro:Adaptive parsing can relocate elements after a website changes, using similarity algorithms
- Pro:Spider framework includes concurrency limits, pause/resume checkpoints, streaming, AutoThrottle, proxy rotation and robots.txt compliance
- Pro:Several fetchers cover plain HTTP with TLS fingerprint impersonation, dynamic browser loading and stealth browsing for Cloudflare Turnstile/Interstitial
- Pro:Includes AI integrations: MCP server, Agent Skill, and page.markdown() output for LLM-ready Markdown
Cons
- Con:Adaptive relocation must be enabled explicitly, with adaptive = True, auto_save=True and then adaptive=True
- Con:The documented anti-bot bypass covers Cloudflare; the README points to a separate sponsor service (Hyper Solutions) for Akamai, DataDome, Kasada and Incapsula
Images
