Rackspace logo

AWS Cloud Engineer III

Rackspace  ·  India, Remote
Remote Full-time Not specified Engineering

Job Description

Role Overview

The L3 DevOps / Cloud Engineer will act as a senior technical resource responsible for leading complex infrastructure and DevOps activities, driving automation, defining technical standards, and improving platform reliability.

The role requires strong hands-on expertise in AWS, Terraform, AWX/Ansible, ArgoCD, GitOps, Kubernetes, CI/CD, and Disaster Recovery.

Key Responsibilities

  • Act as theL3 escalation pointfor complex Cloud, DevOps, Kubernetes, infrastructure, and deployment issues.
  • Lead major technical activities across production and non-production environments.
  • Take ownership of complex infrastructure changes, upgrades, migrations, and platform improvements.
  • Lead and coordinateDisaster Recovery (DR)activities, including:
    • DR planning
    • Restore testing
    • Failover/failback activities
    • Application and infrastructure recovery
    • DR validation
    • Runbook preparation and improvement
    • Design and implement infrastructure usingTerraform.
    • Develop and maintain reusableTerraform modules.
    • Review Terraform code and infrastructure changes implemented by L1/L2 engineers.
    • Define Terraform coding standards, repository structure, and implementation best practices.
    • Strong hands-on experience withAWX / Ansible.
    • Develop reusable Ansible roles, playbooks, and automation workflows.
  • Use AWX to automate:
    • Server configuration
    • Patching
    • Application deployment
    • Infrastructure operations
    • Repetitive BAU activities
    • Identify manual operational activities and convert them into automated workflows.
    • Design and maintainArgoCD-based deployments.
    • Implement and supportGitOps practicesfor Kubernetes and application deployments.
    • Define GitOps repository structures and deployment standards.
  • Manage and troubleshoot ArgoCD:
    • Applications
    • Sync issues
    • Configuration drift
    • Deployment failures
    • Environment promotion
    • Rollback activities
    • Design and maintain Kubernetes workloads running onAmazon EKS.
    • Develop and maintain reusableHelm Charts.
    • Define standards for Helm values, templates, and environment-specific configurations.
    • Design and improve CI/CD deployment processes.
    • Support and improve Jenkins and other CI/CD automation.
    • Define deployment strategies and operational standards.
    • Lead automation initiatives across Cloud and DevOps platforms.
  • Develop automation using:
    • Terraform
    • Ansible / AWX
    • Bash
    • Python
    • CI/CD pipelines
  • Define and enforce technical standards for:
    • Infrastructure as Code
    • Terraform
    • GitOps
    • Kubernetes
    • Helm
    • CI/CD
    • Automation
    • Cloud operations
    • Patching
    • Deployment processes
    • Establish reusable templates, modules, pipelines, and automation frameworks.
    • Perform technical reviews for changes implemented by L1 and L2 engineers.
    • Lead complex production incidents and perform Root Cause Analysis.
    • Identify recurring issues and implement permanent automated solutions.
    • Improve monitoring, logging, alerting, and operational reliability.
    • Participate in architecture and technical design discussions.
    • Lead production deployments, infrastructure upgrades, patching, and maintenance activities.
    • Mentor and provide technical guidance to L1 and L2 engineers.
  • Create and maintain:
    • SOPs
    • Runbooks
    • Technical standards
    • Architecture documentation
    • DR documentation
    • Operational procedures

Apply Now

You'll be redirected to the company's application page

Requirements

  • Amazon EKS
  • Kubernetes
  • Terraform
  • Terragrunt
  • Bash / Python
  • Infrastructure Automation
  • Disaster Recovery
  • Monitoring & Logging
  • Production Troubleshooting
  • Root Cause Analysis
  • Leading technical activities independently
  • Owning complex Cloud and DevOps changes
  • Designing and implementing automation
  • Leading Disaster Recovery activities
  • Implementing GitOps using ArgoCD
  • Building automation using AWX/Ansible
  • Developing and reviewing Terraform code
  • Defining technical standards and best practices
  • Driving operational improvements
  • Mentoring L1/L2 engineers
  • Handling complex production incidents and RCA