Web Scraper for Document Collection
Lokasi klien: 🇬🇧 United Kingdom
Tentang pekerjaan
Need a freelancer to scrape documents from WhstDoTheyKnow, primarily PDFs and some office files, with the potential to collect between 500,000 and 900,000 documents. The work includes building a reliable scraping process, handling large volumes of files, and ensuring consistent data collection. Experience with similar large-scale scraping projects is important.
Please share relevant examples of past work and your approach to managing high-volume document extraction.
The documents must be stored in a file system that records where they came from.
So folder = body name. All documents for that body go into that folder. So we know what documents belong to what body.
We need a super fast turn around time.