5–10 years Data Engineering
- Strong Scala programming
- Strong Apache Spark
- Strong Apache Kafka
- Spark Structured Streaming
- Advanced SQL
- ETL/ELT & Data Pipelines
- Distributed Systems
- Cloud/Databricks exposure
- Required Technical Skills1. Apache Spark – Mandatory
Strong Hands-on Experience With:
- Apache Spark
- Spark Core
- Spark SQL
- Spark DataFrame
- Spark Dataset
- Spark Structured Streaming
- Spark transformations and actions
- RDD concepts
- Joins and aggregations
- Window functions
- Partitioning and repartitioning
- Caching and persistence
- Broadcast joins
- Handling data skew
- Spark job optimization
- Spark cluster execution and troubleshooting
2. Scala – Mandatory
- Strong programming experience in Scala.
- Functional programming concepts.
- Collections and higher-order functions.
- Case classes, traits, objects, pattern matching.
- Exception handling and reusable code development.
- Development of Spark applications using Scala.
- Ability to write clean, modular, scalable, and maintainable Scala code.
3. Apache Kafka – Mandatory
Strong Practical Experience With:
- Kafka architecture
- Kafka brokers
- Topics and partitions
- Producers and consumers
- Consumer groups
- Offsets and offset management
- Replication factor
- Partition strategy
- Message retention
- Kafka Producer/Consumer APIs
- Kafka Streams / Kafka Connect
- Schema Registry
- Avro / JSON serialization
- Consumer lag monitoring
- Kafka troubleshooting
- Kafka performance tuning
- Integration of Kafka with Spark Structured Streaming
4. SQL – Mandatory
Strong SQL Skills Including:
- Complex SQL queries
- Joins
- Subqueries
- CTEs
- Window functions
- Aggregations
- Query optimization
- Data validation and reconciliation
- Stored procedures/functions where applicable
- Relational database concepts
5. Data Engineering
Strong Understanding Of:
- ETL / ELT
- Batch processing
- Real-time/streaming processing
- Data pipelines
- Data lakes
- Data warehouses
- Data modeling
- Dimensional modeling
- Distributed systems
- Data partitioning
- Data quality
- Data governance
- Large-scale data processing
Cloud / Big Data Experience
Experience With At Least One Cloud Platform Is Preferred:
- Required Technical Skills1. Apache Spark – Mandatory
Strong Hands-on Experience With:
- Apache Spark
- Spark Core
- Spark SQL
- Spark DataFrame
- Spark Dataset
- Spark Structured Streaming
- Spark transformations and actions
- RDD concepts
- Joins and aggregations
- Window functions
- Partitioning and repartitioning
- Caching and persistence
- Broadcast joins
- Handling data skew
- Spark job optimization
- Spark cluster execution and troubleshooting
2. Scala – Mandatory
- Strong programming experience in Scala.
- Functional programming concepts.
- Collections and higher-order functions.
- Case classes, traits, objects, pattern matching.
- Exception handling and reusable code development.
- Development of Spark applications using Scala.
- Ability to write clean, modular, scalable, and maintainable Scala code.
3. Apache Kafka – Mandatory
Strong Practical Experience With:
- Kafka architecture
- Kafka brokers
- Topics and partitions
- Producers and consumers
- Consumer groups
- Offsets and offset management
- Replication factor
- Partition strategy
- Message retention
- Kafka Producer/Consumer APIs
- Kafka Streams / Kafka Connect
- Schema Registry
- Avro / JSON serialization
- Consumer lag monitoring
- Kafka troubleshooting
- Kafka performance tuning
- Integration of Kafka with Spark Structured Streaming
4. SQL – Mandatory
Strong SQL Skills Including:
- Complex SQL queries
- Joins
- Subqueries
- CTEs
- Window functions
- Aggregations
- Query optimization
- Data validation and reconciliation
- Stored procedures/functions where applicable
- Relational database concepts
5. Data Engineering
Strong Understanding Of:
- ETL / ELT
- Batch processing
- Real-time/streaming processing
- Data pipelines
- Data lakes
- Data warehouses
- Data modeling
- Dimensional modeling
- Distributed systems
- Data partitioning
- Data quality
- Data governance
- Large-scale data processing
Cloud / Big Data Experience
Strong Understanding Of:
- ETL / ELT
- Batch processing
- Real-time/streaming processing
- Data pipelines
- Data lakes
- Data warehouses
- Data modeling
- Dimensional modeling
- Distributed systems
- Data partitioning
- Data quality
- Data governance
- Large-scale data processing
Cloud / Big Data Experience
- Required Technical Skills1. Apache Spark – Mandatory
Strong Hands-on Experience With:
- Apache Spark
- Spark Core
- Spark SQL
- Spark DataFrame
- Spark Dataset
- Spark Structured Streaming
- Spark transformations and actions
- RDD concepts
- Joins and aggregations
- Window functions
- Partitioning and repartitioning
- Caching and persistence
- Broadcast joins
- Handling data skew
- Spark job optimization
- Spark cluster execution and troubleshooting
- Scala – Mandatory
- Strong programming experience in Scala.
- Functional programming concepts.
- Collections and higher-order functions.
- Case classes, traits, objects, pattern matching.
- Exception handling and reusable code development.
- Development of Spark applications using Scala.
- Ability to write clean, modular, scalable, and maintainable Scala code.
- Apache Kafka – Mandatory
Strong Practical Experience With:
- Kafka architecture
- Kafka brokers
- Topics and partitions
- Producers and consumers
- Consumer groups
- Offsets and offset management
- Replication factor
- Partition strategy
- Message retention
- Kafka Producer/Consumer APIs
- Kafka Streams / Kafka Connect
- Schema Registry
- Avro / JSON serialization
- Consumer lag monitoring
- Kafka troubleshooting
- Kafka performance tuning
- Integration of Kafka with Spark Structured Streaming
- SQL – Mandatory
Strong SQL Skills Including:
- Complex SQL queries
- Joins
- Subqueries
- CTEs
- Window functions
- Aggregations
- Query optimization
- Data validation and reconciliation
- Stored procedures/functions where applicable
- Relational database concepts
- Data Engineering
Strong Understanding Of:
- ETL / ELT
- Batch processing
- Real-time/streaming processing
- Data pipelines
- Data lakes
- Data warehouses
- Data modeling
- Dimensional modeling
- Distributed systems
- Data partitioning
- Data quality
- Data governance
- Large-scale data processing
Cloud / Big Data Experience
Experience With At Least One Cloud Platform Is Preferred:
Experience with at least one cloud platform is preferred:
Skills: kafka,scala,spark,data