
Search by job, company or skills

Job Summary
The Specialist, Cloud Automation & Application Operations is responsible for providing technical leadership and operational oversight across cloud platforms, automation services, application operations, DevOps, and Site Reliability Engineering (SRE) functions. The role ensures cloud and application services are reliable, secure, scalable, and highly automated through the implementation of Infrastructure as Code (IaC), CI/CD pipelines, observability, and operational best practices. The Technical Lead serves as the primary technical escalation point for complex operational issues, platform reliability challenges, service incidents, and automation initiatives while mentoring engineers and driving continuous improvement.
Job Responsibilities:
1. Cloud Operations Technical Leadership
• Manage cloud platforms to ensure secure, scalable, and reliable service delivery. Monitor cloud resources and coordinate operational activities to maintain platform availability and support business-critical services.
• Optimize cloud utilization and implement operational standards to improve efficiency, reliability, and consistency.
• Lead incident resolution efforts to minimize service disruptions and maintain operational stability.
2. Automation & Platform Engineering
• Develop automation solutions and implement Infrastructure as Code (IaC) practices to improve operational efficiency and deployment consistency.
• Automate provisioning and configuration processes to accelerate service delivery and reduce manual effort.
• Standardize operational workflows and maintain automation frameworks to enhance service quality, scalability, and reliability.
• Drive platform engineering initiatives to improve agility, operational resilience, and continuous improvement.
3. Application Operations Management
• Manage application platforms to ensure service availability and business continuity.
• Monitor application health and coordinate support activities to maintain performance and minimize service disruptions.
• Resolve operational issues within agreed service levels and maintain operational procedures to ensure consistent support delivery.
• Drive continuous improvements to enhance application reliability, stability, and user experience.
4. DevOps & CI/CD Operations
• Manage CI/CD platforms and automate deployment pipelines to deliver software reliably and efficiently.
• Implement DevOps practices to strengthen collaboration and accelerate releases. Govern release activities and improve deployment processes to increase success rates and reduce operational risks.
• Support application modernization and cloud-native adoption to enhance agility, scalability, and innovation.
5. Site Reliability Engineering (SRE)
• Define service reliability standards and establish service level objectives (SLOs) to ensure consistent operational performance and measurable service outcomes.
• Monitor service health indicators to proactively detect and address issues before they impact users.
• Reduce operational toil through automation and engineering improvements to increase efficiency and scalability.
• Improve platform resilience and drive continuous service improvement through performance analysis to strengthen business continuity and service reliability.
6. Monitoring, Incident, Problem & Service Operations
• Coordinate major incident response and manage technical escalations to restore services quickly and minimize business impact.
• Perform root cause analysis and drive problem resolution initiatives to prevent recurring issues and improve operational stability.
• Review service performance trends and implement proactive improvements to enhance service reliability, reduce MTTD and MTTR, and strengthen overall operational effectiveness.
• Implement monitoring and observability solutions to improve service visibility and operational awareness across cloud and application environments.
• Develop dashboards and reporting capabilities to support data-driven decision-making and performance management.
• Support AIOps and predictive monitoring initiatives to enhance service reliability, operational efficiency, and proactive issue detection.
7. Technical Governance & Stakeholder Management
• Provide technical leadership and mentorship to cloud and operations engineers.
• Review solution designs, operational readiness assessments, and deployment plans.
• Collaborate with development, infrastructure, cybersecurity, and business teams. Coordinate with vendors and cloud service providers on operational matters.
• Develop operational standards, documentation, and best practices. Support audits, compliance requirements, and governance activities.
Qualification & Work Experience
1. Master's Degree in Telecommunications, Computer Engineering, Computer Science, Software Engineering, Electronics Engineering, Data Analytics, or related field from a recognized university.
2. MBA or equivalent is an added advantage
3. Minimum 8 to 10 years in Technical Lead position or Applications Developer or Systems/Applications Management.
4. Proven experience in:
a. Large-scale datacenter operations and infrastructure management
b. Multi-cloud and hybrid cloud orchestration
c. AI-driven automation / AIOps implementation
d. Managing automation platforms
e. Applications management and support
f. Vendor and partner management
g. Experience supporting mission-critical production environments with 24x7 operational requirements.
h. Hands-on experience in automation, CI/CD, cloud platforms, monitoring, and incident management.
5. Strong exposure to ITIL, DevOps, and cybersecurity frameworks.
6. Relevant certifications (preferred):
a. Microsoft Azure Administrator / Azure Solutions Architect
b. AWS Solutions Architect Associate/Professional
c. Google Professional Cloud Architect
d. Certified Kubernetes Administrator (CKA)
e. HashiCorp Terraform Associate
f. Red Hat Ansible Automation
g. DevOps Institute Certifications
h. ITIL Foundation or ITIL Managing Professional
Technical Competencies/Skills
• Site Reliability Engineering (SRE), Systems Administration, or DataOps environment.
• Production experience operating workloads in at least one major cloud provider (AWS, Azure, or GCP)
• Ops & Orchestration like Kubernetes, Docker, Helm, Linux Systems Administration
• Automation & IaC like Terraform (state management), Ansible, Bash, Python
• CI/CD & Deployment like GitHub Actions, GitLab CI, Jenkins, ArgoCD
• Monitoring & Logging as Datadog, Prometheus, Grafana, ELK Stack, Splunk
• Data Operations as Apache Airflow scheduling, SQL tuning, Snowflake/Databricks operations
• Cloud orchestration (multi-cloud, hybrid cloud)
• AI-driven operations (AIOps, predictive analytics)
• High availability & disaster recovery planning
• Infrastructure automation & DevOps practices
• Vendor and technology management
Behavioral Competencies/Skills
• Technical leadership and vision
• Strong decision-making under pressure
• Stakeholder and executive influence
• Crisis and incident leadership
• Innovation and transformation mindset
• Results-driven and accountability-focused
• Strong communication (technical & non-technical)
• Team leadership and talent development
• Governance and risk awareness
• Continuous improvement mindset
Candidate must be willing to work in Cyberjaya and Malaysian Citizen.
Job ID: 151351621
Skills:
Data Center Infrastructure, Microsoft Azure, Data Center Operations, AWS, NOC Operations, Itil, Cloud Operations, Google Cloud, Devops, Appdynamics, Dynatrace, Splunk, Cybersecurity Operations, Infrastructure Monitoring, Certified Data Centre Professional, Facilities Management Monitoring, Facilities Operations and Monitoring, Cybersecurity frameworks, Enterprise monitoring and observability platforms, Infrastructure Operations Management, SOC Operations, Cloud Monitoring Services, Automation AIOps, SIEM Security Analytics, Observability Platforms, Service Management Functions, Certified Data Centre Operations Professional
Skills:
Ip Routing, switching, cloud-based networking, network security platforms, IP network operations, automation scripting tools, virtualization platforms, NFV infrastructure
Skills:
Outlook, Ms Office, MS Teams, Substation Automation, Sharepoint, Sales Skills, Communication
Skills:
Control M, Batch Monitoring
Skills:
tax accounting , automation, Transfer Pricing, Direct Tax, Indirect Tax, Tax Risk Management, Tax Compliance, AI-enabled workflows