Python/PostgreSQL Developer: Data Discovery, Search & MCP MVP
Lokasi klien: 🇦🇺 Australia
Tentang pekerjaan
We need a hands-on Python/PostgreSQL developer to turn existing directory data into a useful, searchable product, with repeatable refresh, net-new discovery and secure read-only MCP access. We want practical delivery and sensible reuse, not a long enterprise transformation programme.
EXISTING ASSETS
An earlier inventory identified approximately 6.58 million source rows across 24 inputs, covering podcasts, journalists/media, YouTube, Instagram, TikTok, investors and blogs/SEO. These are source rows, not unique or verified profiles. Much dates to 2023, with some 2025 material. Existing scripts, consolidated data and a basic search interface may be reusable. The product has no final name.
FIRST RELEASE: BUILD SOMETHING WE CAN USE
Start with podcasts and journalists to prove the reusable system. Propose a precise, affordable launch population and field list. Profile the wider repository, but do not assume every row or category needs deep enrichment before launch.
1. Briefly inspect representative files and existing code; agree what to reuse and the fixed implementation scope.
2. Import an agreed launch subset into PostgreSQL, using neutral internal IDs, normalised fields and conservative deduplication. Keep uncertain matches separate.
3. Refresh commercially important fields: activity, current organisation/role, website and appropriate professional contact routes. Treat old information as a historical baseline, not as newly verified. Unknown stays unknown.
4. Demonstrate repeatable discovery of current entities missing from the relevant historical population, not merely missing from the small pilot sample. Identify duplicates/rebrands separately from genuinely new records.
5. Provide a simple authenticated search/filter/profile/export interface and documented retrieval API. Reuse existing components where suitable; no bespoke design exercise.
6. Expose that same retrieval layer through a small authenticated read-only MCP server, with bounded results, usage controls and setup documentation. No arbitrary SQL access or outbound sending.
7. Deliver rerunnable code, scheduled refresh/discovery for the initial categories, error handling, cost reporting, basic tests and a tested backup/restore procedure in client-owned accounts.
QUALITY AND BOUNDARIES
A usable contact route matters more than a filled email column. Separate syntax/deliverability, affiliation and contact purpose. Do not infer permission to contact from an email check. Keep fresh observations and check dates distinguishable from historical imports. Use permitted APIs/RSS/structured sources before costly browser automation; disclose access, licensing and coverage limitations. No bypassing platform access controls. No paid services without approval.
The final product must be brand-neutral and independent of historical vendor IDs/CDNs. Any irreversible retirement of legacy migration material is a separately approved post-acceptance task, subject to retention requirements; do not delete originals during this initial engagement.
BUDGET AND DELIVERY
The displayed US$2,000 is an indicative target for this lean first release, not a budget to deeply refresh all 6.58M rows or deliver every future category. Submit your best realistic fixed-price proposal, including lower-priced options through reuse or reduced scope. If more is required, explain exactly why and offer the smallest useful alternative. Do not underbid and leave essential work unspecified.
We need an indicative TOTAL through a working first release, not only an audit fee. A short paid inspection may confirm the final scope; include it in your milestone breakdown rather than proposing an open-ended discovery programme. Work is funded only after a contract is agreed. Any substantial trial is paid.
Propose your fastest credible schedule, weekly working demonstrations and clear acceptance criteria. Separately price expansion to the remaining categories/full population, customer billing or a polished frontend if excluded, and expected monthly infrastructure/acquisition/maintenance. Continuous new discovery and contact enrichment at full scale are not assumed included in a small launch quote.
WHO SHOULD APPLY
One accountable, hands-on technical owner with real Python, SQL/PostgreSQL, web acquisition/API/RSS, entity-resolution and backend experience. MCP experience is useful, but data quality and reliable delivery come first. A small team is welcome if you name the engineer doing the work and disclose subcontracting. Worldwide applicants welcome. Communication and demonstrations must remain on Upwork.
Show one comparable system you personally built, especially how it found NEW entities and handled duplicates/changes. Existing reusable code is welcome where you have the rights and disclose its dependencies. No unpaid custom implementation or speculative sales promises required. Please answer the screening questions with concrete examples, an itemised estimate, availability and the main risks you would resolve first.