Site Reliability Engineering Lead
Job Description
• Leadership & Strategy• Define and implement SRE best practices across the organization.
• Proven expertise in production support, engineering, disaster recovery (DCR), automation, and cloud operations
• Mentor and guide a team of SREs, fostering growth
• Collaborate with senior stakeholders to align reliability goals with business objectives.
• Reliability & Performance• Establish SLIs, SLOs, and SLAs for critical services and ensure adherence.
• Drive initiatives to improve system and reduce operational toil.
• Excellent in designing systems that detect and remediate issues without manual intervention – Self Healing systems, Runbook automation
• Exposure to tools like Gremlin, Chaos Monkey, AWS FIS to simulate outages and improve fault tolerance
• Incident Management• Act as the primary point of escalation for critical production issues and lead major incident response, root cause analysis, and postmortems.
• Perform detailed post-incident investigations to identify underlying causes. Document findings and share learnings to prevent recurrence.
• Implement preventive measures and continuous improvement processes.
• Observability• Champion monitoring, logging, and alerting strategies using tools like Prometheus, Grafana, ELK, and AWS CloudWatch.
• Build real-time dashboards to visualize system health and reliability metrics.
• Configure intelligent alerting based on anomaly detection and thresholds.
• Combine metrics, logs, and traces to enable root cause analysis and reduce Mean Time to Resolution (MTTR).
• Knowledge of AIOps or ML-based anomaly detection for proactive reliability management.
• Collaboration• Work closely with development teams to integrate reliability into application design and deployment
• Promote a culture of shared responsibility for uptime and performance across engineering teams.
Requirements
Department: Technology
Function: Information Technology
Experience Level: Not Applicable