We're Hiring: Observability Architect
Location: Kuala Lumpur, Malaysia
Employment Type: Full-time
Experience: 8+ Years Relevant experience
Are you passionate about building enterprise-scale observability platforms that improve reliability, reduce operational complexity, and enable AIOps-driven operations We are looking for an experienced Observability Architect to lead the strategy, architecture, and governance of our enterprise observability ecosystem across multi-cloud and hybrid environments.
In this role, you will define and implement end-to-end observability capabilities across cloud platforms, applications, infrastructure, Kubernetes, and enterprise services. You will play a key role in driving operational excellence by leveraging Dynatrace, ServiceNow ITOM, and modern observability technologies to deliver intelligent monitoring, event correlation, root cause analysis, and self-healing automation.
What You'll Do
- Define and govern enterprise observability architecture, standards, and best practices.
- Design and implement monitoring, logging, tracing, APM, Digital Experience Monitoring, and service health dashboards.
- Lead the enterprise architecture and governance of the Dynatrace platform.
- Architect integrations between Dynatrace and ServiceNow ITOM for event management, service mapping, incident automation, and operational intelligence.
- Design AIOps capabilities including anomaly detection, event correlation, predictive alerting, and automated remediation.
- Develop observability patterns for Kubernetes, cloud-native applications, APIs, and microservices.
- Define enterprise logging strategies using Elastic, OpenSearch, and cloud-native logging services.
- Establish telemetry standards using Prometheus, Grafana, and OpenTelemetry.
- Define monitoring standards across Azure and Alibaba Cloud services.
- Collaborate with Cloud, DevOps, Security, Infrastructure, and Application teams to embed observability into platform and solution design.
- Mentor engineering teams and promote SRE and operational excellence practices across the organization.
Required Skills & Experience
- 12+ years of experience in Infrastructure, Cloud, Platform Engineering, SRE, or Enterprise Architecture.
- Proven experience designing enterprise observability solutions.
- Strong expertise with Dynatrace (architecture, governance, dashboards, RUM, Synthetic Monitoring, Distributed Tracing, Davis AI).
- Hands-on experience integrating Dynatrace with ServiceNow ITOM/Event Management.
- Strong experience with Azure Monitor, Log Analytics, Application Insights, Azure Managed Prometheus, and Azure Managed Grafana.
- Experience with Kubernetes (AKS) and cloud-native observability.
- Knowledge of Prometheus, Grafana, OpenTelemetry, Elastic, or OpenSearch.
- Experience implementing AIOps, automation, self-healing, and event correlation.
- Strong understanding of SLIs, SLOs, error budgets, monitoring governance, and service reliability engineering.
Nice to Have
- Alibaba Cloud observability (CloudMonitor, ARMS, Log Service, ACK).
- Infrastructure as Code (Terraform, Ansible).
- CI/CD and DevOps practices.
- Insurance or BFSI industry experience.
- Enterprise architecture and platform governance experience.
Why Join Us
- Lead enterprise-wide observability transformation.
- Work with modern cloud-native and AIOps technologies.
- Drive architecture decisions across a large-scale multi-cloud environment.
- Collaborate with high-performing engineering, cloud, security, and operations teams.
- Make a measurable impact on platform reliability, automation, and operational excellence.
If you're excited about designing the future of enterprise observability and enabling resilient, intelligent operations, we'd love to hear from you. Apply now or reach out directly for a confidential discussion.