Remote

Web Scraper for Document Collection

Lokasi klien: 🇬🇧 United Kingdom

Tentang pekerjaan

Need a freelancer to scrape documents from WhstDoTheyKnow, primarily PDFs and some office files, with the potential to collect between 500,000 and 900,000 documents. The work includes building a reliable scraping process, handling large volumes of files, and ensuring consistent data collection. Experience with similar large-scale scraping projects is important.

Please share relevant examples of past work and your approach to managing high-volume document extraction.

The documents must be stored in a file system that records where they came from.

So folder = body name. All documents for that body go into that folder. So we know what documents belong to what body.

We need a super fast turn around time.