[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes
Job Description
Project – the aim you'll have
We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work.
Position – how you’ll contribute
• Support the deployment, operation, and reliability of production services running on Kubernetes.
• Monitor service health and investigate production incidents across distributed applications.
• Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
• Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
• Support CI/CD, GitOps-based deployments, observability, and production monitoring.
• Work within a client-directed backlog and established priorities.
Requirements
Department: DevOps, Cloud & Infrastructure
Function: Information Technology
Experience Level: Mid-Senior Level