Scraping Engineer
Clera · 1 day ago
About the Role
You'll be an embedded member of a fast-moving AI startup team, owning the reliability and quality of a web extraction infrastructure that processes millions of pages daily through hundreds of scripts. Reporting to a Forward Deployed Engineer, this role is critical to keeping a high-volume data pipeline running smoothly and accurately.
What You'll Do
-
Maintain and monitor a triage queue, resolving broken scrapers and data quality alerts to keep operations running smoothly.
-
Build and deploy new web scraping scripts for websites requiring data extraction.
-
Validate scraped data for accuracy and investigate discrepancies.
-
Create dashboards to visualize and monitor scraped data in real time.
-
Collaborate with AI agents to fill capability gaps, fix issues, and improve existing scripts.
What We're Looking For
-
1+ years of hands-on, professional experience building or maintaining web scraping solutions.
-
Proficiency with TypeScript and Node.js for building and debugging scraping scripts.
-
Proficiency with SQL for data querying and validation.
-
Experience with Puppeteer or similar browser automation libraries.
-
Experience debugging and fixing broken scraping scripts in production environments.
-
Familiarity with message queues (e.g. RabbitMQ) and asynchronous job processing systems.
-
Experience with Redis or similar in-memory caching systems.
-
Experience with Google Cloud Platform or equivalent cloud infrastructure.
-
Comfort with bash scripting, git, and gRPC.
-
Exposure to BigQuery or similar data warehousing solutions is a plus.
-
Experience building dashboards or data visualization tools is a plus.
-
Strong async communicator who flags blockers proactively and thrives with ambiguity in a fast-paced environment.
-
Availability for 8+ hours per day with overlap during US business hours.
Location
Fully remote — all time zones welcome, provided you can maintain overlap with US business hours.
Originally posted on Himalayas