Jemmy · Twelve scraping scenarios

Explore a fictional Capri travel catalogue and collect its data with Jemmy or your preferred tool. Each case starts an isolated run that lasts 30 minutes.

How it works and fair use

S01 · HTML catalogue

HTML listings and detail pages

S02 · Independent JSON pages

Pagination and concurrent requests

S03 · Opaque cursors

A chain of dependent pages

S04 · Partial GraphQL responses

Variables, cursors and a recoverable field error

S05 · JavaScript interface

The same catalogue through a rendered DOM and an API

S06 · Protobuf enrichment

Search, detail pages and a binary data join

S07 · Expiring sessions

Login form, CSRF, session cookie and reauthentication

S08 · Token refresh and permissions

Two accounts, short-lived access tokens and single-use refresh

S09 · Account quotas

Sliding-window quotas, concurrency limits and Retry-After

S10 · JavaScript challenge

A nonce, SHA-256 and a temporary clearance cookie

S11 · TLS fingerprint

Observed ClientHello and matching request headers

S12 · Authenticated recovery

Client interruption, a scheduled 503 and session renewal