All projects

Washington DC Business Registry Paginator

Single-CLI Python scraper that bulk-pages the DLCP paginationadvance-search API.

The challenge

DLCP offers the registry only through a paginated search UI - no CSV, no 'download all' button.

The solution

A single-CLI Python scraper (washingtondc_pagination.py) that calls DLCP's paginationadvance-search API directly with exact headers (x-app-route, referer), % wildcard business-name search, token-based pagination, minimal record extraction, and tagged timestamped logs. --start-page enables mid-run resume.

Results

DC DLCP registry

Target

--start-page

Resume

Streaming JSON

Output