Data Platform Engineer (102-08SENG-01)
OpsBrasil Serviços Cloud LTDA · 7 hours ago
This role converts a large Azure Data Factory estate into Databricks workflows on AWS. The scope for year one: 2,094 ADF pipelines to migrate — built as reusable templates rather than one-by-one — 9,859 pipeline activities to translate (some map directly, others need rewriting as Lambda or Step Functions), 471 Spark dataflows to move onto Databricks on AWS, and 4 Databricks workspaces (Dev, QA, Pre-prod, Prod) to rehost, including notebook paths and Unity Catalog rewiring. This is a regulated environment, so reconciling migrated data against source systems is part of the definition of done, not an afterthought.
Requirements
What you will do
-
Convert Azure Data Factory pipelines into Databricks workflows on AWS, building reusable templates rather than migrating one at a time.
-
Rehost Databricks workspaces onto AWS and migrate ADLS Gen2 storage to S3.
-
Rewrite ADF Web Activities as Lambda functions or Step Functions tasks, and replace ADF-specific scaling with native Databricks mechanisms.
-
Build and tune PySpark transformations for production data volumes.
-
Replace Azure Synapse Serverless with Databricks SQL Warehouse.
-
Reconcile migrated data against source systems as part of the definition of done.
Required
-
Production experience with Databricks: workspaces, jobs and workflows. The central skill for this role.
-
Strong Spark and PySpark experience for real data volumes, including tuning.
-
Production-grade Python.
-
Experience building or migrating Azure Data Factory pipelines, with a solid understanding of the ADF activity model.
-
AWS data services: S3, Glue, Athena, Lambda and Step Functions.
-
Advanced SQL, including reading and reasoning about stored procedures.
-
Professional written and spoken English.
Nice to have
Delta Lake, Unity Catalog, Azure Synapse, Terraform, Airflow/MWAA, dbt, Kafka, Databricks certification, data modeling, Great Expectations, SAS/analytics platform integration, CRM data.
Engagement details
-
Full-time
-
100% remote
-
Open to candidate from all LATAM
Highlights
Databricks, PySpark, Python, Azure Data Factory, AWS (S3, Glue, Athena, Lambda, Step Functions), SQL, Delta Lake, Unity Catalog, Azure Synapse, Terraform, Airflow/MWAA, dbt, Kafka
Originally posted on Himalayas