

Search by job, company or skills

Position Objective:
The Observability Architect is responsible for defining, designing, implementing, and governing enterprise observability capabilities across AIA multi-cloud, hybrid, and containerized environments. The role will establish consistent monitoring, logging, tracing, event correlation, service health, and operational intelligence practices across Azure, Alibaba Cloud, and related enterprise platforms.
Dynatrace will be the primary observability platform, integrated with ServiceNow ITOM to support event management, incident enrichment, service mapping, root cause analysis, and operational automation. The role will also define complementary patterns for logging and open-source monitoring platforms such as Elastic, OpenSearch, Prometheus, Grafana, and OpenTelemetry.
A key objective of this role is to drive AIOps-enabled operations and self-healing automation to reduce alert noise, improve Mean Time To Detect (MTTD), reduce Mean Time To Resolve (MTTR), and improve overall platform and application reliability.
Roles and Responsibilities:
Observability Architecture & Design
• Define and maintain enterprise observability reference architecture, standards, patterns, and governance across AIA Group and Business Units.
• Design end-to-end observability for infrastructure monitoring, application performance monitoring, distributed tracing, logging, digital experience monitoring, network observability, and service health dashboards.
• Establish observability KPIs, SLOs, SLIs, error budgets, alerting principles, and service reliability reporting standards.
• Review solution designs and ensure observability requirements are embedded from architecture and delivery stages.
Dynatrace Platform Architecture
• Lead architecture and governance of Dynatrace as the enterprise observability plat
form.
• Design monitoring standards for applications, Kubernetes, virtual machines, databases, middleware, APIs, network services, and cloud-native services.
• Define Dynatrace tagging standards, management zones, dashboard patterns, service mapping, synthetic monitoring, real user monitoring, and Davis AI adoption.
• Drive platform configuration, onboarding patterns, operational dashboards, reporting, and observability data retention standards.
ServiceNow ITOM Integration
• Architect integration between Dynatrace and ServiceNow ITOM Event Management / ITOM Health capabilities.
• Design event correlation, alert enrichment, topology-aware service impact analysis, and automated incident creation workflows.
• Define noise reduction, deduplication, priority mapping, escalation, and operational ownership models.
• Support integration of observability insights into ITSM processes, command center dashboards, and operational reporting.
AIOps & Auto-Healing Automation
• Define and implement AIOps operating patterns for anomaly detection, predictive alerting, root cause identification, and closed-loop remediation.
• Design auto-healing workflows integrated with Dynatrace, ServiceNow ITOM, automation platforms, and cloud-native services.
• Identify repeatable operational failure patterns and convert them into automated remediation runbooks where appropriate.
• Drive continuous improvement to reduce manual intervention, alert fatigue, incident recurrence, MTTD, and MTTR.
Logging & Observability Data Platforms
• Architect centralized logging and log analytics solutions using Elastic, OpenSearch, cloud-native log services, and enterprise logging standards.
• Define log collection, normalization, enrichment, indexing, retention, access control, and cost governance standards.
• Establish patterns for correlation across logs, metrics, traces, events, configuration, and service topology.
• Ensure logging platforms support operational troubleshooting, security visibility, auditability, and compliance requirements.
Open Source Monitoring & Cloud Native Observability
• Design monitoring and visualization solutions using Prometheus, Grafana, OpenTelemetry, and related open-source monitoring platforms.
• Define observability patterns for AKS, Alibaba Cloud ACK, containers, microservices, APIs, service mesh, and DevOps pipelines.
• Establish metrics federation, dashboarding, alerting, and integration approaches with Dynatrace and enterprise platforms.
• Promote consistent instrumentation and telemetry standards across modern application architectures.
Azure & Alibaba Cloud Native Monitoring
• Define monitoring standards for Azure platform services including Azure Monitor, Log Analytics, Application Insights, Network Watcher, Azure Managed Prometheus, Azure Managed Grafana, Azure Advisor, and Azure Resource Health.
• Define monitoring standards for Alibaba Cloud services including CloudMonitor, Log Service (SLS), ActionTrail, ARMS, Managed Service for Prometheus, Security Center, and related observability and alarm services.
• Ensure cloud-native monitoring capabilities are integrated with enterprise observability, ITOM, incident, and reporting processes.
• Drive cost-aware telemetry design across Azure and Alibaba Cloud environments.
Security, Governance & Compliance
• Embed security-by-design, least-privilege access, data protection, and compliance requirements into observability architecture.
• Support audits, risk assessments, regulatory reviews, and evidence requirements related to monitoring, logging, and service reliability.
• Define guardrails for observability platform access, retention, data classification, dashboard sharing, and operational reporting.
• Maintain observability documentation, runbooks, standards, design patterns, and operational controls.
Stakeholder & Technical Leadership
• Act as the enterprise subject matter expert for observability, monitoring, logging, AIOps, and self-healing automation.
• Collaborate with Cloud Architecture, Cloud Engineering, Operations, Security, DevOps, Application, Service Management, and Business Unit teams.
• Mentor engineering and operations teams on observability best practices, platform onboarding, dashboarding, and incident reduction techniques.
• Drive observability transformation initiatives and promote SRE-aligned operational practices across the organization.
Minimum Job Requirements:
Skills:
• 10+ years relevant experience in Cloud Architecture, Infrastructure, Operations, Observability, or Application Performance Management.
• 5+ years practical experience designing and governing enterprise observability platforms.
• Strong hands-on experience with Dynatrace, including APM, infrastructure monitoring, Kubernetes monitoring, dashboards, management zones, service mapping, and Davis AI
capabilities.
• Experience integrating observability platforms with ServiceNow ITOM / Event Management / IT
SM processes.
• Strong experience with logging platforms such as Elastic Stack, OpenSearch, or equivalent enterprise log analytics solutions.
• Experience with open-source monitoring and visualization platforms such as Prometheus, Grafana, and OpenTelemetry.
• Hands-on experience with Azure native monitoring services including Azure Monitor, Log Analytics, Application Insights, Network Watcher, Azure Managed Prometheus, and Azure Man
aged Grafana.
• Hands-on experience with Alibaba Cloud native monitoring services including CloudMonitor, Log Service (SLS), ActionTrail, ARMS, Managed Service for Prometheus, and related observability services.
• Strong understanding of Kubernetes, containers, microservices, APIs, network monitoring, distributed tracing, cloud networking, IAM, and security principles.
• Experience implementing AIOps, event correlation, anomaly detection, automated remediation, and self-healing operations.
• Experience working in enterprise or regulated environments is highly desirable.
• Relevant professional certifications such as Dynatrace Associate/Professional, Microsoft Azure Solutions Architect Expert, Alibaba Cloud Professional Architect, CKA, ITIL, SRE Foundation, or TOGAF will be an advantage.
• Sound understanding of IT partner ecosystem and partner collaboration in a multinational
corporation.
• Experience in top-tier multinational corporation will be an advantage.
• Strong problem-solving and analytical skills.
• Excellent communication, stakeholder management, and collaboration skills.
• Ability to translate complex operational telemetry into actionable service reliability insights.
• Ability to operationalize disruptive technology services, including building implementation roadmaps for observability, AIOps, and self-healing automation.
Job ID: 151552811