
Search by job, company or skills

Job Summary
The Specialist, Cloud & Datacentre Monitoring Operations is responsible for leading the technical operations, monitoring governance, and operational assurance of the organization's cloud and data centre environments. The role provides technical leadership for the Network Operations Centre (NOC) and Facilities Monitoring Operations, ensuring proactive monitoring, rapid incident response, service reliability, and operational resilience across critical infrastructure platforms. The position serves as the primary technical escalation point for monitoring-related events, major incidents, and operational issues affecting network infrastructure, cloud platforms, data centre services, and mission-critical facilities. The role drives observability, automation, event management, and continuous service improvement to maintain high availability and operational excellence.
Job Responsibilities:
1. Technical Leadership and Monitoring Governance
• Lead enterprise monitoring platforms and command centre operations to ensure reliable service visibility and operational effectiveness.
• Establish monitoring standards and governance to maintain consistent operations.
• Provide technical leadership and escalation support to resolve complex issues efficiently.
• Drive monitoring optimization and automation to improve operational performance and resilience.
2. Network Operation Centre (NOC)
• Technical lead of 24x7 monitoring of network, cloud, data centre, and infrastructure services to ensure service availability and operational resilience.
• Oversee real-time monitoring, proactive event detection, and rapid service restoration to minimize business impact.
• Manage escalation and major incident processes to accelerate resolution and maintain service stability.
• Implement observability and predictive monitoring capabilities to improve visibility, operational efficiency, and performance.
3. Facilities Monitoring Operation
• Technical lead of facilities monitoring to ensure the availability and reliability of mission-critical data centre infrastructure.
• Oversee power, cooling, environmental, security, and fire protection systems to minimize operational risks and service disruptions.
• Coordinate incident response, maintenance activities, and vendor support to accelerate issue resolution and improve facility resilience.
• Integrate monitoring capabilities to enhance visibility, operational efficiency, and proactive risk management.
4. Incident and Problem Management
• Coordinate incident and problem management process to ensure rapid service restoration, minimize business disruption, and improve operational stability.
• Coordinate cross-functional teams, manage major incidents, and communicate incident status effectively to accelerate resolution and maintain stakeholder confidence.
• Monitor incident trends, conduct post-incident reviews, and perform root cause analysis to identify recurring issues and implement permanent corrective actions.
• Analyse operational risks, drive preventive measures, and promote continuous service improvement to enhance service reliability, operational resilience, and organizational learning.
5. Change Management
• Coordinate change management governance across infrastructure, cloud, and application services to ensure controlled and compliant service delivery.
• Coordinate change assessment, approval, testing, and implementation to minimize operational risks and service disruptions.
• Coordinating CAB activities, monitor change performance, and drive process improvements to enhance change success rates.
• Balance service stability with business agility while ensuring compliance with audit and regulatory requirements.
6. Automation Monitoring & Service Assurance
• Lead enterprise observability and monitoring strategies to improve visibility, operational resilience, and service performance across infrastructure, applications, cloud, facilities, and security domains.
• Integrate monitoring platforms into a centralized command centre, implement AIOps and automation capabilities to accelerate issue detection and response, and reduce alert fatigue through intelligent event correlation.
• Enable predictive analytics, real-time dashboards, and executive reporting to support proactive decision-making and continuous operational improvement.
7. Capacity & Service Performance Management
• Monitor infrastructure utilization and service performance to identify capacity risks and support informed decision-making.
• Recommend proactive improvements to enhance service availability, operational readiness, and platform resilience.
• Produce operational insights and performance reports to support capacity planning and continuous service optimization.
8. Stakeholder & Vendor Management
• Collaborate with internal teams, vendors, and service providers to ensure effective monitoring operations and timely issue resolution.
• Review monitoring requirements for new services and infrastructure to maintain operational readiness and service visibility.
• Provide technical recommendations to improve monitoring capabilities and support audit, compliance, and operational assurance requirements.
Qualification & Work Experience
1. Master's Degree in Telecommunications, Computer Engineering, Computer Science, Software Engineering, Electronics Engineering, Data Analytics, or related field from a recognized university.
2. MBA or equivalent is an added advantage
3. Minimum 7 to 10 years in Technical Lead position or Service Operations Center Management.
4. Proven experience in:
a. NOC Operations
b. SOC Operations
c. Data Center Operations
d. Service Management Functions
e. Cloud Monitoring Services
f. Facilities Management & Monitoring
5. Experience managing 24x7 mission-critical environments.
6. Experience implementing enterprise monitoring and observability platforms.
7. Strong exposure to ITIL, DevOps, and cybersecurity frameworks.
8. Relevant certifications (preferred): ITIL Managing Professional / ITIL Expert, b. Cisco CCNP / CCIE, c. VMware VCP / VCA
9. Microsoft Azure Administrator / Architect
10. AWS Certified Solutions Architect
11. Google Cloud Professional Cloud Architect/Operations
12. Certified Data Centre Professional (CDCP)
13. Certified Data Centre Operations Professional (CDCOP)
14. Splunk Certified Administrator or Dynatrace Associate Certification or AppDynamics Certification (or any)
Technical Competencies/Skills
• Infrastructure Operations Management
• Cloud Operations
• Data Center Infrastructure
• Facilities Operations and Monitoring
• Infrastructure Monitoring
• Cybersecurity Operations
• SIEM & Security Analytics
• Observability Platforms
• Automation & AIOps
Behavioral Competencies/Skills
• Technical leadership and vision
• Strong decision-making under pressure
• Stakeholder and executive influence
• Crisis and incident leadership
• Innovation and transformation mindset
• Results-driven and accountability-focused
• Strong communication (technical & non-technical)
• Team leadership and talent development
• Governance and risk awareness
• Continuous improvement mindset
Candidate must be willing to work in Cyberjaya and Malaysian Citizen.
Job ID: 151372003

Skills:
Devops, Cloud Orchestration, Automation, Monitoring
Skills:
Ip Routing, switching, cloud-based networking, network security platforms, IP network operations, automation scripting tools, virtualization platforms, NFV infrastructure
Skills:
Metasploit, Ethical Hacking, Splunk, Burp Suite, Linux System Administration, Penetration Testing, Network Protocols, Security Controls, Information Security, Web Application Security, Qualys, Cybersecurity, Owasp Top 10, Iam, SIEM platforms, Nessus, AWS Cloud services
Skills:
Mean Stack, Node.js, Openshift, Kubernetes, Docker, Devops Tools, Jenkins, MongoDB, Git Actions, ITIL Practices
Skills:
Gcp, Secure Code Review, Azure, Penetration Testing, AWS, scanning tools for mobile API and web application security testing, design of highly-available and highly-secure solutions, threat modelling, design of container-based infrastructures, secure design review, development of mobile applications