Search by job, company or skills

Custom Software Engineer

Early Applicant
  • Posted 2 days ago
  • Be among the first 10 applicants

Job Description

Project Role : Custom Software Engineer

Project Role Description : Lead the effort to design, build and configure applications, acting as the primary point of contact.

Must have skills : PySpark

Good to have skills : Python (Programming Language)

Minimum 5 Year(s) Of Experience Is Required

Educational Qualification : 15 years full time education

Summary:

We are seeking an experienced Senior Data Engineer to design, develop, and optimize enterprise-scale data products capable of processing high-volume, complex datasets. The ideal candidate has deep expertise in PySpark, distributed data processing, and performance optimization, with a strong understanding of modern Lakehouse architectures and cloud-native data platforms.

Roles & Responsibilities:

Design, develop, and maintain scalable data products using PySpark and Spark SQL.

Build reusable, modular, and production-ready data pipelines supporting enterprise analytics and AI use cases.

Develop data models for structured, semi-structured, and streaming data.

Performance Engineering

Optimize complex PySpark transformations and Spark SQL queries processing billions of records.

Improve application performance through:

Efficient partitioning strategies

Data skew mitigation

Broadcast joins

Bucketing and sorting

Caching and persistence

Predicate pushdown

Adaptive Query Execution (AQE)

File compaction and optimization

Analyze Spark execution plans and identify performance bottlenecks.

Optimize memory utilization, shuffle operations, executor configuration, and cluster resource consumption.

Large-Scale Data Processing

Build pipelines capable of handling TB to PB-scale datasets.

Process batch and near real-time data efficiently while maintaining SLA commitments.

Ensure scalability, resiliency, and fault tolerance of distributed workloads.

Data Quality & Reliability

Implement automated data validation and reconciliation checks.

Develop monitoring, alerting, and logging for production pipelines.

Perform root cause analysis for production failures and implement preventive improvements.

Cloud & Lakehouse Engineering

Develop solutions on cloud-based data platforms.

Work with Delta Lake, Iceberg, or Parquet-based architectures.

Optimize storage layout, partitioning, and file management for improved query performance.

Collaboration

Partner with architects, product owners, analysts, and data scientists to translate business requirements into scalable data products.

Participate in design reviews, code reviews, and architecture discussions.

Mentor junior engineers on PySpark best practices and performance tuning.

Professional & Technical Skills:

Python

PySpark

Spark SQL

SQL (Advanced)

Additional Information:

Big Data

Apache Spark

Distributed data processing

Data partitioning

Spark optimization

Shuffle optimization

Memory tuning

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152103667

Similar Jobs

Bengaluru, India

Skills:

software design patterns Version Control SystemsApi DevelopmentTesting FrameworksDatabase integrationDebugging techniquesCollaborative development workflowsPython Programming LanguageObject-oriented programming

Beware of Scammers

We don’t charge money for job offers