Rackspace logo

AWS Cloud Engineer IV

Rackspace  ·  India, Remote
Remote Full-time Not specified Engineering

Job Description

Job Profile Summary

The L4 Cloud DevOps Engineer serves as a technical subject matter expert (SME) for cloud architecture and operational engineering, focusing on AWS, Terraform (IaC), operating systems, Kubernetes (EKS), and modern DevOps practices. The role designs, implements, and supports cloud infrastructure and deployment automation, leads incident response and DR activities, mentors junior engineers, and collaborates with stakeholders to deliver secure, reliable, and scalable solutions that align with business objectives.

Career Level Summary

  • Requires deep technical expertise across AWS, Terraform, OS administration, Kubernetes, and DevOps tooling.
  • Leads others to solve complex technical problems and provides technical leadership on projects.
  • Works independently on sophisticated engineering tasks; escalates or seeks guidance for only the most complex, ambiguous situations.
  • May provide functional leadership and mentorship to L1/L2 engineers.

Key Responsibilities

  • Act as the L3 escalation point and technical SME for complex AWS, OS, Kubernetes, infrastructure and deployment issues.
  • Design and implement infrastructure as code using Terraform (and optionally CloudFormation), building reusable modules and enforcing coding standards.
  • Implement and maintain CI/CD pipelines and GitOps practices (ArgoCD), including pipeline design, environment promotion, rollback, and drift remediation.
  • Develop automation for configuration management and operational tasks using Ansible/AWX, scripting (Bash, PowerShell, Python), and other automation tools.
  • Design, deploy and maintain Kubernetes workloads (Amazon EKS), Helm charts, and environment-specific templating.
  • Perform Linux and Windows administration: hardening, patching, performance tuning, and automated deployments.
  • Lead incident response and major production incidents; perform RCA and implement permanent automated fixes.
  • Lead and coordinate Disaster Recovery planning and testing, including runbooks, failover/failback, and validation.
  • Define and enforce technical standards for IaC, Kubernetes, CI/CD, monitoring, logging, and operational processes.
  • Improve observability: monitoring, alerting, logging and operational reliability across platforms.
  • Mentor and provide technical guidance to L1 and L2 engineers; review and approve infrastructure and deployment changes.
  • Create and maintain SOPs, runbooks, architecture documentation, and DR documentation.
  • Collaborate with product and customer engineering teams to deliver cloud-based solutions, migrations, and optimizations.

Experience

10+ years of relevant experience in cloud engineering, infrastructure, DevOps, or a related field is required.

L4 / SME Expectations

  • Lead and own complex engineering changes and transformations across cloud platforms.
  • Drive automation, standardization, and reliability improvements in operations and delivery.
  • Define technical standards, review peer code, and enforce best engineering practices.
  • Act as a technical escalation point for production incidents and lead root cause analysis.
  • Provide mentorship and hands-on guidance to junior engineers, elevating team capability.

Certifications

  • Preferred: AWS Certifications (Solutions Architect, DevOps Engineer Professional).
  • Beneficial: CNCF/Kubernetes Certifications (CKA, CKAD, CKS), Azure/GCP certifications as applicable.

Apply Now

You'll be redirected to the company's application page

Requirements

  • Strong communication and stakeholder management; able to translate technical concepts for non-technical audiences
  • Influence and coaching: mentor team members, promote best practices, and drive continuous improvement
  • Analytical problem solving: diagnosing complex production issues and identifying root cause and long-term fixes
  • Collaboration: work across disciplines to design, implement, and operate cloud solutions
  • Project and time management: prioritize work, track deliverables, and meet deadlines while balancing multiple initiatives
  • Technical documentation: create clear runbooks, SOPs, architecture diagrams, and standards
  • Expert-level knowledge of AWS services, architecture patterns, security, networking, and operational best practices
  • Deep understanding of Infrastructure as Code concepts and Terraform module design, state management and security
  • Extensive knowledge of Linux and Windows internals, administration, automation, and security hardening
  • In-depth Kubernetes architecture and operations experience (preferably Amazon EKS), including Helm and GitOps patterns
  • Familiarity with common security and compliance frameworks (e.g., NIST, HIPAA, PCI) and secure operational controls
  • Cloud & IaC: Expert in AWS and Terraform (including Terragrunt patterns), reusable module development, state management, and IaC security
  • Kubernetes & Containers: Design, operate and troubleshoot EKS clusters, Helm chart development, container best practices, and GitOps (ArgoCD)
  • CI/CD & Automation: Build and maintain CI/CD pipelines, artifact management, automated testing, and deployment automation (Jenkins, GitHub Actions, GitLab CI, etc.)
  • Configuration Management: Use Ansible/AWX for system configuration, patching, application deployment and repetitive operational tasks
  • Operating Systems: Expert-level Linux administration and solid Windows server administration; automation via scripting (Bash, PowerShell, Python)
  • Version Control: Strong Git skills (branching strategies, code review, hooks, access controls)
  • Monitoring & Observability: Implementing monitoring, logging, and alerting (CloudWatch, Prometheus, ELK/EFK or similar)
  • Security & IAM: Implement secure IAM practices, secret management, network security and encryption in cloud environments
  • Disaster Recovery & Resilience: DR planning, regular testing, and runbook creation for reliable failover and recovery