Company Overview
We are a talent acquisition and staff augmentation firm, and we are currently hiring on behalf of one of our clients, an organization in the retail and e-commerce technology space. Our client is looking to onboard a Site Reliability Engineer (SRE) L2 across multiple locations in India for a remote or hybrid role that can be offered on a full-time, part-time, or contractual basis, depending on the candidate’s preference and eligibility.
Job Summary
We are looking for an experienced Site Reliability Engineer (SRE) L2 with strong expertise in Microsoft Azure, Dynatrace, and application monitoring to support business-critical cloud applications. The ideal candidate has hands-on experience with Azure PaaS services, observability platforms, incident response, root cause analysis, and production support. This role calls for proactive monitoring, hands-on troubleshooting, and a strong focus on keeping cloud-hosted applications highly available and reliable.
Key Responsibilities
• Monitor, maintain, and provide production support for applications hosted on the Microsoft Azure platform, ensuring high availability and operational stability.
• Diagnose and resolve application and infrastructure incidents through comprehensive end-to-end troubleshooting to minimize service disruptions.
• Analyze production alerts and investigate performance issues using Azure Monitor, Application Insights, Log Analytics, and Dynatrace to identify root causes and implement timely resolutions.
• Analyze application logs and telemetry using Kusto Query Language (KQL) and Dynatrace Query Language (DQL).
• Monitor and troubleshoot Azure API Management (APIM), Azure Functions, Service Bus, and other Azure-native services.
• Find out the originating reason of application failures by tracking requests across APIs, Azure Functions, messaging services, and backend components.
• Configure and manage alerts, dashboards, and monitoring rules to enable proactive incident detection.
• Use Dynatrace features such as Smartscape, Problems & Events, Distributed Tracing, Synthetic Monitoring, and Davis AI for performance analysis.
• Participate in production on-call rotations and provide support for P1/P2 incidents.
• Lead technical troubleshooting during major incidents and collaborate with development and infrastructure teams for timely resolution.
• Prepare detailed Root Cause Analysis (RCA) reports and recommend preventive measures.
• Work closely with DevOps, Application Development, and Cloud Infrastructure teams to improve platform reliability and observability.
• Support continuous service improvements through automation and operational excellence initiatives.
Required Skills
Microsoft Azure
• Azure Monitor, Application Insights, Log Analytics
• Kusto Query Language (KQL)
• Azure API Management (APIM), Azure Functions, Azure Service Bus
• Azure Alerts & Action Groups, Azure Portal
Monitoring & Observability
• Dynatrace, including Problems & Events Feed, Smartscape, Distributed Tracing, Synthetic Monitoring, and Dynatrace Query Language (DQL)
• Alternative observability tools such as New Relic or Datadog are also acceptable
Incident Management
• Production support, P1/P2 incident handling, and major incident management
• Perform Root Cause Analysis (RCA) and drive problem management initiatives to resolve underlying issues and improve service reliability.
• SLA management and on-call support
AI Experience
• Hands-on experience using Dynatrace Davis AI for automated anomaly detection and performance analysis.
• Comfort using AI-assisted tools for log analysis, alert triage, or incident summarization is a plus.
Managerial Experience
• Experience leading technical troubleshooting during major incidents and coordinating across development and infrastructure teams.
• Comfortable collaborating with DevOps, Application Development, and Cloud Infrastructure teams to drive reliability improvements.
Operational Experience
• 5–7 years of experience in production support, incident management, and cloud application reliability.
• Experience participating in on-call rotations and managing P1/P2 incidents under time pressure.
• Track record of preparing RCA reports and driving preventive measures to reduce recurring incidents.
Qualifications
• Any graduate or postgraduate degree.
• 5–7 years of relevant experience in SRE, production support, or application monitoring roles.
Certifications
• Microsoft Certified: Azure Administrator Associate or Azure Solutions Architect certification (preferred, not mandatory).
• Dynatrace certification (preferred, not mandatory).
Submit Resume at dp@digitalxnode.com
Explore More Job Opportunities at digitalxnode.com/fte/