Multiple sub‑roles are available under this posting: General Data Engineer, Cloudera‑Hadoop‑NiFi ETL Engineer, and Azure Databricks / ADF Data Engineer. Candidates with hands‑on experience in any one of these technology stacks are welcome to apply. Final role allocation will be determined during the interview phase based on each candidate's technical background.
Key Responsibilities
- Develop and migrate batch and streaming ETL/ELT pipelines using Databricks, PySpark, Spark SQL and Delta Lake.
- Convert legacy data-processing logic into scalable lakehouse patterns while preserving business rules and data quality.
- Optimise Spark jobs, partitioning, storage layouts, caching, joins and cluster usage for performance and cost efficiency.
- Implement data ingestion, transformation, validation, error handling, logging and observability.
- Support source-to-target mapping, migration execution, reconciliation and production cutover activities.
- Create reusable frameworks, coding standards, automated tests and CI/CD pipelines for data engineering.
- Troubleshoot data, performance and pipeline failures across development, test and production environments.
- Collaborate with architects, testers, governance specialists and platform operations teams.
Required Skills & Experience
- Practical experience with Databricks, PySpark, Spark SQL and Delta Lake.
- Strong SQL and data-transformation skills with good understanding of data modelling and file/table formats.
- Experience building production-grade ETL/ELT pipelines and integrating multiple source systems.
- Knowledge of performance tuning, distributed processing and large-volume data handling.
- Experience with Git, CI/CD and software engineering practices for data platforms.
- Ability to diagnose data-quality, transformation and performance issues.
- Exposure to banking data domains or regulated environments is preferred.
Preferred / Advantageous
- Azure Databricks, ADF and ADLS experience.
- Kafka or other streaming technology experience.
- Python software-engineering experience beyond notebook development.
- Cloudera / Hadoop / NiFi ETL Engineer
Key Responsibilities
- Analyse Cloudera/Hadoop workloads, HDFS data, Hive/Impala objects, Spark jobs, Kafka flows and NiFi pipelines.
- Document dependencies, schedules, interfaces, transformations, security settings and operational requirements.
- Modify, refactor or convert legacy ETL and ingestion flows for migration to the target platform.
- Support data extraction, movement, validation, reconciliation and cutover activities.
- Troubleshoot Hadoop, NiFi, Kafka, Hive and Spark workload issues during migration and parallel-run phases.
- Assist with compatibility analysis and identification of unsupported or replacement components.
- Develop migration utilities, scripts and technical runbooks where required.
- Work closely with Databricks engineers and architects on source-to-target workload mapping.
Required Skills & Experience
- Hands-on experience in Cloudera CDH/CDP or comparable Hadoop distributions.
- Good knowledge of HDFS, Hive, Spark, Kafka and NiFi.
- Strong Linux, SQL and scripting skills; Python or Shell experience is desirable.
- Experience developing or supporting ETL/data-ingestion pipelines in enterprise environments.
- Ability to perform dependency analysis, troubleshooting and performance diagnostics.
- Understanding of migration validation, reconciliation and cutover processes.
- Experience in banking or large-scale regulated data platforms is preferred.
Preferred / Advantageous
- Cloudera certification or equivalent platform experience.
- Experience migrating workloads from Cloudera to Databricks or other cloud lakehouse platforms.
- Knowledge of Kerberos, Ranger/Sentry, Atlas or similar security/governance components.
- Azure Data Engineer - Databricks / ADF / ADLS
Key Responsibilities
- Design and develop data pipelines using Azure Data Factory, Azure Databricks and Azure Data Lake Storage.
- Implement secure ingestion and integration from on-premises and cloud sources into the target data platform.
- Develop orchestration, scheduling, monitoring, retry and error-handling patterns for production pipelines.
- Integrate Databricks with Azure services, identity, networking, Key Vault and enterprise security controls.
- Create CI/CD processes for notebooks, jobs, infrastructure configuration and pipeline deployments.
- Optimise pipeline performance, storage usage, compute resources and operational reliability.
- Support migration testing, reconciliation, cutover and production stabilisation.
- Document technical designs, operating procedures and support runbooks.
Required Skills & Experience
- Hands-on experience with Azure Databricks, ADF and ADLS.
- Strong SQL and Python/PySpark skills.
- Experience with Azure identity, networking, security and secret-management concepts.
- Knowledge of data integration, orchestration, file formats, partitioning and lakehouse design.
- Experience with Git-based CI/CD and deployment automation.
- Ability to troubleshoot cloud data pipelines and performance issues.
- Experience with enterprise or banking data workloads is preferred.
Preferred / Advantageous
- Azure Data Engineer or Databricks certification.
- Experience with Synapse, Event Hubs, Functions or related Azure data services.
- Infrastructure-as-code experience such as Terraform or Bicep.