Role Overview Role Overview: Looking for a Site Reliability Engineer (SRE) to support the reliability, availability, and performance of business‑critical systems. This role focuses on AWS cloud infrastructure, DevOps tools, and core SRE practices. You will work closely with development, platform, and operations teams to ensure systems are stable, scalable, and well monitored. Key Responsibilities: Reliability & Operations: Support high availability, scalability, and performance of production systems Implement and maintain SLIs, SLOs, and SLAs for services Identify and reduce operational toil through automation and process improvement Support design and implementation of fault‑tolerant and resilient systems AWS & Cloud Operations: Manage and operate systems hosted on AWS (EC2, EKS/ECS, RDS, S3, Lambda, CloudWatch, IAM, VPC) Support cloud deployments and infrastructure changes following best practices Assist with backup, disaster recovery, and resiliency planning DevOps & Automation: Work with CI/CD pipelines and DevOps tools to support reliable deployments Use Infrastructure as Code tools such as Terraform or CloudFormation Automate repetitive tasks using scripts (Python, Bash, etc.) Incident Management: Participating in production incident response, troubleshooting, and service restoration Perform root cause analysis (RCA) and contribute to post‑incident reviews Help implement …
Jobs Finder matches jobs to your profile and auto-writes an ATS-friendly CV + cover letter for each, delivered to your Telegram.
Get my tailored CV →