SKILL·CDE025

headless-web-scraping

Name: headless-web-scraping
Author: pjt222

pjt222

업데이트됨 1 month ago

8 조회

디자인apiautomationdesigndata

정보

이 스킬은 scrapling 파이썬 라이브러리를 활용해 강력한 웹 스크래핑을 가능하게 하며, 사이트의 방어 수준에 따라 HTTP, 스텔스 크로미움, 완전한 브라우저 자동화 중 적절한 방법을 자동으로 선택합니다. JavaScript 렌더링 페이지, 봇 방지 보호 기능, 구조화된 데이터 추출을 위한 복잡한 DOM 탐색까지 처리합니다. 동적 콘텐츠나 고급 차단 메커니즘으로 인해 WebFetch가 실패할 때 사용하세요.

빠른 설치

Claude Code

문서

Headless Web Scraping

Extract data from web pages that resist simple HTTP requests — JS-rendered content, Cloudflare-protected sites, and dynamic SPAs — using scrapling's three-tier fetcher architecture and CSS-based data extraction.

When to Use

Target page requires JavaScript rendering (SPA, React, Vue)
Site has anti-bot protections (Cloudflare Turnstile, TLS fingerprinting)
You need structured extraction of multiple elements via CSS selectors
Simple WebFetch or requests.get() returns empty or blocked responses
Extracting tabular data, link lists, or repeated DOM structures at scale

Inputs

Required: Target URL or list of URLs to scrape
Required: Data to extract (CSS selectors, field names, or description of target elements)
Optional: Fetcher tier override (default: auto-select based on site behavior)
Optional: Output format (default: JSON; alternatives: CSV, Python dict)
Optional: Rate limit delay in seconds (default: 1)

Procedure

Step 1: Select Fetcher Tier

Determine which scrapling fetcher matches the target site's defenses.

# Decision matrix:
# 1. Fetcher        — static HTML, no JS, no anti-bot (fastest)
# 2. StealthyFetcher — Cloudflare/Turnstile, TLS fingerprint checks
# 3. DynamicFetcher  — JS-rendered SPAs, click/scroll interactions

# Quick probe: try Fetcher first, escalate on failure
from scrapling import Fetcher

fetcher = Fetcher()
response = fetcher.get("https://example.com/target-page")

if response.status == 200 and response.get_all_text():
    print("Fetcher tier sufficient")
else:
    print("Escalate to StealthyFetcher or DynamicFetcher")

Signal	Recommended Tier
Static HTML, no protection	`Fetcher`
403/503, Cloudflare challenge page	`StealthyFetcher`
Page loads but content area is empty	`DynamicFetcher`
Need to click buttons or scroll	`DynamicFetcher`
altcha CAPTCHA present	None (cannot be automated)

Got: One of the three tiers is identified. For most modern sites, StealthyFetcher is the correct starting point.

If fail: If all three tiers return blocked responses, check whether the site uses altcha CAPTCHA (proof-of-work challenge that cannot be bypassed). If so, document the limitation and provide manual extraction instructions instead.

Step 2: Configure the Fetcher

Set up the selected fetcher with appropriate options.

from scrapling import Fetcher, StealthyFetcher, DynamicFetcher

# Tier 1: Fast HTTP with TLS fingerprint impersonation
fetcher = Fetcher()
fetcher.configure(
    timeout=30,
    retries=3,
    follow_redirects=True
)

# Tier 2: Headless Chromium with anti-detection
fetcher = StealthyFetcher()
fetcher.configure(
    headless=True,
    timeout=60,
    network_idle=True  # wait for all network requests to settle
)

# Tier 3: Full browser automation
fetcher = DynamicFetcher()
fetcher.configure(
    headless=True,
    timeout=90,
    network_idle=True,
    wait_selector="div.results"  # wait for specific element before extracting
)

Got: Fetcher instance is configured and ready. No errors on instantiation. For StealthyFetcher and DynamicFetcher, a Chromium binary is available (scrapling manages this automatically on first run).

If fail:

playwright or browser binary not found -- run python -m playwright install chromium
Timeout on configure() -- increase timeout value or check network connectivity
Import error -- install scrapling: pip install scrapling

Step 3: Fetch and Extract Data

Navigate to the target URL and extract structured data using CSS selectors.

# Fetch the page
response = fetcher.get("https://example.com/target-page")

# Single element extraction
title = response.find("h1.page-title")
if title:
    print(title.get_all_text())

# Multiple elements
items = response.find_all("div.result-item")
for item in items:
    name = item.find("span.name")
    price = item.find("span.price")
    print(f"{name.get_all_text()}: {price.get_all_text()}")

# Get attribute values
links = response.find_all("a.product-link")
urls = [link.get("href") for link in links]

# Get raw HTML content of an element
detail_html = response.find("div.description").html_content

Key API reference:

Method	Purpose
`response.find("selector")`	First matching element
`response.find_all("selector")`	All matching elements
`element.get("attr")`	Attribute value (href, src, data-*)
`element.get_all_text()`	All text content, recursively
`element.html_content`	Raw inner HTML

Got: Extracted data matches the visible page content. Elements are non-None and text content is non-empty for populated pages.

If fail:

find() returns None -- inspect the actual HTML (response.html_content) to verify the selector; the page may use different class names than expected
Empty text from get_all_text() -- content may be inside shadow DOM or an iframe; try DynamicFetcher with a wait_selector
Do NOT use .css_first() -- this is not part of the scrapling API (common confusion with other libraries)

Step 4: Handle Failures and Edge Cases

Implement fallback logic for CAPTCHA detection, empty responses, and session requirements.

import time

def scrape_with_fallback(url, selector):
    """Try each fetcher tier in order, with CAPTCHA detection."""
    tiers = [
        ("Fetcher", Fetcher),
        ("StealthyFetcher", StealthyFetcher),
        ("DynamicFetcher", DynamicFetcher),
    ]

    for tier_name, tier_class in tiers:
        fetcher = tier_class()
        fetcher.configure(headless=True, timeout=60)

        try:
            response = fetcher.get(url)
        except Exception as error:
            print(f"{tier_name} failed: {error}")
            continue

        # Detect CAPTCHA / challenge pages
        page_text = response.get_all_text().lower()
        if "altcha" in page_text or "proof of work" in page_text:
            print(f"altcha CAPTCHA detected -- cannot automate")
            return None

        if response.status == 403 or response.status == 503:
            print(f"{tier_name} blocked (HTTP {response.status}), escalating")
            continue

        result = response.find(selector)
        if result and result.get_all_text().strip():
            return result.get_all_text()

        print(f"{tier_name} returned empty content, escalating")

    print("All tiers exhausted. Manual extraction required.")
    return None

Got: Function returns extracted text on success, or None with a diagnostic message when all tiers fail. CAPTCHA pages are detected and reported rather than retried indefinitely.

If fail:

All tiers return 403 -- the site blocks all automated access (common with WIPO, TMview, some government databases); document the URL as requiring manual access
Timeout errors -- the page may be behind a slow CDN; increase timeout to 120s
Session/cookie errors -- the site may require login; add cookie handling or authenticate first

Step 5: Rate Limiting and Ethical Scraping

Implement delays and respect site policies before running at scale.

import time
import urllib.robotparser

def check_robots_txt(base_url, target_path):
    """Check if scraping is allowed by robots.txt."""
    rp = urllib.robotparser.RobotFileParser()
    rp.set_url(f"{base_url}/robots.txt")
    rp.read()
    return rp.can_fetch("*", f"{base_url}{target_path}")

def scrape_urls(urls, selector, delay=1.0):
    """Scrape multiple URLs with rate limiting."""
    results = []
    fetcher = StealthyFetcher()
    fetcher.configure(headless=True, timeout=60)

    for url in urls:
        response = fetcher.get(url)
        data = response.find(selector)
        if data:
            results.append(data.get_all_text())

        time.sleep(delay)  # respect the server

    return results

Ethical scraping checklist:

Check robots.txt before scraping -- respect Disallow directives
Use a minimum 1-second delay between requests
Identify your scraper with a descriptive User-Agent when possible
Do not scrape personal data without legal basis
Cache responses locally to avoid redundant requests
Stop immediately if you receive a 429 (Too Many Requests)

Got: Scraping runs at a controlled rate. robots.txt is checked before bulk operations. No 429 responses are triggered.

If fail:

429 Too Many Requests -- increase delay to 3-5 seconds, or stop and retry later
robots.txt disallows the path -- respect the directive; do not override it
IP ban -- stop scraping immediately; the rate limiting was insufficient. If access is legitimate (public data, ToS-permitted, robots.txt-respected) and you must continue, see rotate-scraping-proxies for network-layer escalation

Validation

Correct fetcher tier is selected (not over- or under-powered for the target)
configure() method is used (not deprecated constructor kwargs)
CSS selectors match actual page structure (verified against page source)
.find() / .find_all() API is used (not .css_first() or other library methods)
CAPTCHA detection is in place (altcha pages are reported, not retried)
Rate limiting is implemented for multi-URL scraping
robots.txt is checked before bulk operations
Extracted data is non-empty and structurally correct

Pitfalls

Using .css_first() instead of .find(): scrapling uses .find() and .find_all() for element selection -- .css_first() belongs to a different library and will raise AttributeError
Starting with DynamicFetcher: Try Fetcher first, then escalate -- DynamicFetcher is 10-50x slower due to full browser startup
Constructor kwargs instead of configure(): scrapling v0.4.x deprecated passing options to the constructor; use the configure() method
Ignoring altcha CAPTCHA: No fetcher tier can solve altcha proof-of-work challenges -- detect them early and fall back to manual instructions
No rate limiting: Even if the site does not return 429, aggressive scraping can get your IP banned or cause service degradation
Assuming stable selectors: Website CSS classes change frequently -- validate selectors against current page source before each scraping campaign

Related Skills

rotate-scraping-proxies -- network-layer escalation when client-side stealth is exhausted and IP bans block legitimate, ToS-permitted access
use-graphql-api -- structured API queries when the site offers a GraphQL endpoint (preferred over scraping)
serialize-data-formats -- converting extracted data to JSON, CSV, or other formats
deploy-searxng -- self-hosted search engine that aggregates results from multiple sources
forage-solutions -- broader pattern for gathering information from diverse sources

GitHub 저장소

pjt222/agent-almanac

경로: i18n/caveman-lite/skills/headless-web-scraping

agentsagentskillsai-assisted-developmentclaude-codeskillsteams

FAQ

Frequently asked questions

What is the headless-web-scraping skill?

headless-web-scraping is a Claude Skill by pjt222. Skills package instructions and resources that Claude loads on demand, so Claude can perform headless-web-scraping-related tasks without extra prompting.

How do I install headless-web-scraping?

Use the install commands on this page: add headless-web-scraping to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.

What category does headless-web-scraping belong to?

headless-web-scraping is in the Design category, tagged api, automation, design and data.

Is headless-web-scraping free to use?

Yes. headless-web-scraping is listed on AIMCP and free to install. It runs inside Claude, so no separate service account is required to use the skill itself.

연관 스킬

executing-plans

디자인

executing-plans 스킬은 검토 체크포인트가 포함된 통제된 배치로 실행할 완전한 구현 계획이 있을 때 사용합니다. 이 스킬은 계획을 불러와 비판적으로 검토한 후, 소규모 배치(기본값 3개 작업)로 작업을 실행하면서 각 배치 사이에 진행 상황을 아키텍트 검토를 위해 보고합니다. 이를 통해 내재된 품질 관리 체크포인트를 갖춘 체계적인 구현이 보장됩니다.

스킬 보기

requesting-code-review

디자인

이 스킬은 코드 변경 사항을 요구 사항에 따라 분석하기 위해 코드 리뷰어 하위 에이전트를 호출합니다. 작업 완료 후, 주요 기능 구현 후, 또는 메인 브랜치에 병합하기 전에 사용해야 합니다. 이 리뷰는 현재 구현체와 원래 계획을 비교하여 문제를 조기에 발견하는 데 도움이 됩니다.

스킬 보기

connect-mcp-server

디자인

이 스킬은 개발자들이 HTTP, stdio 또는 SSE 전송 방식을 통해 MCP 서버를 Claude Code에 연결하는 포괄적인 가이드를 제공합니다. GitHub, Notion 및 사용자 정의 API와 같은 외부 서비스를 통합하기 위한 설치, 구성, 인증 및 보안을 다룹니다. MCP 통합 설정, 외부 도구 구성 또는 Claude의 모델 컨텍스트 프로토콜 작업 시 활용하세요.

스킬 보기

web-cli-teleport

디자인

이 스킬은 작업 분석을 기반으로 개발자가 Claude Code 웹 인터페이스와 CLI 인터페이스 중 선택할 수 있도록 돕고, 두 환경 간 원활한 세션 텔레포트를 가능하게 합니다. 웹, CLI 또는 모바일 환경 전환 시 세션 상태와 컨텍스트를 관리하여 워크플로를 최적화합니다. 다양한 단계에서 서로 다른 도구가 필요한 복잡한 프로젝트에 사용하세요.

스킬 보기