Repo Frameworks Limited info

D4Vinci/Scrapling

Python framework for web scraping, from single requests to concurrent crawls, with adaptive element relocation, anti-bot fetchers, spiders, proxy rotation, CLI and MCP server.

  • 86.5k GitHub stars
  • Python
  • ⚖️ BSD-3-Clause
  • 🎯 Intermediate
D4Vinci/Scrapling preview image

What it is

Scrapling is an adaptive web scraping framework that covers everything from a single request to a full-scale crawl. Its parser relocates elements when pages change. Its fetchers can bypass anti-bot systems like Cloudflare Turnstile. Its Scrapy-like spider framework supports concurrent multi-session crawls with pause/resume, proxy rotation and AutoThrottle.

Who it's for

  • Web scrapers who want selectors that survive website design changes
  • Developers building concurrent, resumable crawls in Python
  • Teams feeding AI agents or RAG pipelines with scraped pages, through the MCP server, Agent Skill or Markdown output
  • Scrapy or BeautifulSoup users looking for a familiar API

Requirements

Requirements

  • Python (the README links a supported-versions badge but states no version)
  • A Playwright Chromium or Google Chrome browser for DynamicFetcher browser automation

Examples

Adaptive fetching and selection

python
python
from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher
StealthyFetcher.adaptive = True
p = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)  # Fetch website under the radar!
products = p.css('.product', auto_save=True)                                        # Scrape data that survives website design changes!
products = p.css('.product', adaptive=True)                                         # Later, if the website structure changes, pass `adaptive=True` to find them!

What it does: Fetches a page with StealthyFetcher, saves the matched elements with auto_save=True, and relocates them later with adaptive=True if the site structure changes.

Scaling up to a full crawl with a Spider

python
python
from scrapling.spiders import Spider, Response

class MySpider(Spider):
  name = "demo"
  start_urls = ["https://example.com/"]

  async def parse(self, response: Response):
      for item in response.css('.product'):
          yield {"title": item.css('h2::text').get()}

MySpider().start()

What it does: Defines a Scrapy-like spider with start_urls and an async parse callback that yields items, then starts the crawl.

Pros & cons

Pros

  • Pro:Adaptive parsing can relocate elements after a website changes, using similarity algorithms
  • Pro:Spider framework includes concurrency limits, pause/resume checkpoints, streaming, AutoThrottle, proxy rotation and robots.txt compliance
  • Pro:Several fetchers cover plain HTTP with TLS fingerprint impersonation, dynamic browser loading and stealth browsing for Cloudflare Turnstile/Interstitial
  • Pro:Includes AI integrations: MCP server, Agent Skill, and page.markdown() output for LLM-ready Markdown

Cons

  • Con:Adaptive relocation must be enabled explicitly, with adaptive = True, auto_save=True and then adaptive=True
  • Con:The documented anti-bot bypass covers Cloudflare; the README points to a separate sponsor service (Hyper Solutions) for Akamai, DataDome, Kasada and Incapsula

Images