Mid levelcloud

Site Reliability Engineer
Interview Questions

Covering SRE interview questions — SLIs/SLOs/SLAs, incident management, capacity planning, and chaos engineering.. Free, no signup required.

10 questions ready

Q1
Walk me through how you would design a monitoring and alerting strategy for a microservices application running on Kubernetes in production. What metrics would you prioritize, and what tools have you used?
Why they ask this:* SREs must understand observability fundamentals, cloud-native architecture, and the ability to design systems that catch failures before they impact users. This reveals depth in Prometheus, Grafana, ELK stacks, and alerting best practices.
Q2
Describe your experience with infrastructure-as-code (IaC) tools like Terraform or CloudFormation. How have you used them to manage cloud resources, and what challenges did you encounter with state management or drift detection?
Why they ask this:* Cloud SREs must automate infrastructure provisioning and management. This tests practical experience with IaC, version control for infrastructure, and understanding of reproducible deployments.
Q3
Explain how you would troubleshoot a sudden spike in application latency in a distributed system. Walk me through your debugging methodology and what observability tools you'd leverage.
Why they ask this:* This assesses root-cause analysis skills, understanding of distributed tracing (Jaeger, Zipkin), log aggregation, and the ability to work systematically under pressure—core SRE competencies.
Q4
What is your experience with incident management and post-incident reviews? Describe how you've implemented or improved runbooks, playbooks, or automated remediation for common incidents.
Q5
Tell me about a time when you were on-call and had to respond to a critical production incident. What was the situation, what actions did you take to resolve it, and what did you learn afterward?
Q6
Describe a situation where you had to balance shipping new features quickly with maintaining system reliability. How did you handle conflicting priorities between development and operations teams?
Q7
Give me an example of a system outage or failure you were involved in that you could have prevented with better planning or automation. What would you do differently, and have you implemented that change?
Q8
What would you do if you discovered a critical memory leak in a core production service running on your cloud infrastructure, but the development team is unavailable for 8 hours? How would you mitigate the impact?
Q9
How would you handle a situation where your automated deployment pipeline has a high failure rate, causing deployment delays and frustrating the development team, but you don't have time to fully overhaul it?
Q10
Imagine a cloud vendor announces a critical security vulnerability affecting your infrastructure. You have 48 hours to patch thousands of instances across multiple environments. What's your approach?
🔒

7 questions locked

Upgrade to unlock all 10 questions with answer guides, videos & PDF

Upgrade to unlock →

Want questions tailored to a specific company?

Try the full generator →