Join EPAM Malaysia as an Application Support Engineer, where you will play a key role in sustaining the performance and reliability of mission-critical platforms.
You'll manage containerized workloads, optimize database operations, and troubleshoot complex issues across production environments. Partner closely with engineering, operations, and product teams to ensure seamless service delivery and drive continuous improvement across our application ecosystem. This role is ideal for mid-level engineers seeking to deepen their expertise within a modern, cloud-native stack while contributing to critical production systems.
Responsibilities
- Manage and resolve issues detected by monitoring systems or reported by internal/external users, ensuring timely and high-quality responses
- Monitor application performance, identify degradation patterns and proactively implement corrective actions
- Automate operational workflows and routine tasks through scripting and tooling
- Lead or contribute to root cause analysis (RCA) and postmortems and ensure corrective measures are implemented
- Maintain and enhance project documentation, runbooks and operational guidelines to support platform reliability
- Work closely with development, product, and operations teams to implement solutions and improve platform stability
- Contribute to process improvements that enhance platform scalability, reliability and operational efficiency
Requirements
- 3–5 years of hands-on experience in software application support, platform operations, or production engineering, with demonstrated ability to troubleshoot and optimize production systems
- Proven experience operating PostgreSQL in production environments, including installation, configuration, performance tuning and troubleshooting complex SQL workloads
- Operational experience with Docker and Kubernetes, including workload management, deployment troubleshooting and containerized application optimization
- Solid understanding of Linux administration, including logs, system services, permissions and performance diagnostics
- Practical experience with monitoring and log analysis using tools such as Prometheus, Grafana, ELK, Loki, or similar
- Ability to lead troubleshooting, contribute to RCA and implement durable solutions in high-availability environments
- Strong communication and collaboration skills with the ability to articulate technical details clearly in English (spoken and written)
Nice to have
- Experience with scripting (Shell, Python, or similar) for automation and operational efficiency
- Exposure to cloud-native ecosystems and basic DevOps practices
- Familiarity with CI/CD pipelines, automation tools and infrastructure-as-code frameworks