Vultr logo

Senior System Engineer

Vultr  ·  United States, Chennai
Not specified Full-time Senior Infrastructure

Job Description

Vultr Cares

  • Medical Insurance stipend paid annually

  • 9 Company-Paid Holidays

  • Generous Leave Policy + 1 month paid sabbatical every 5 years + Anniversary Bonus each year

  • Professional Development Reimbursement

  • Internet reimbursement

  • Fitness membership reimbursement

  • Company paid Wellable subscription

Join Vultr

Vultr is expanding its India presence and is building its firstGlobal Integrated Operations Command Center (GIOC) in Chennai– the 24×7 nerve center for monitoring, triage, and first response across Vultr’s globalcloud, systems, network, and security operations.

We are seekingSenior Systems Engineersto serve as the first line of response for Vultr’s systems estate –compute, storage, operating systems, and virtualization.This is a frontline operations role for engineers who thrive in a fast-paced, alert-driven environment and want to grow into deeper SRE, systems, and platform engineering careers. You must be comfortable working a rotational shift model – including nights, weekends, and holidays – to sustain follow-the-sun coverage.

Role Overview

TheSenior Systems Engineeris the entry point of the GIOC incident lifecycle. You monitor health signals across Vultr’s systems estate, acknowledge and triage alerts withindefined first-response SLAs,execute documented runbooks to resolve common issues, and escalate cleanly to Senior Engineers when an issue falls outside L2 scope. The role ensuresfast, accurate first responsethat protects customer experience and keeps incidents moving toward resolution.

You operate from a consolidatedsingle-pane-of-glass dashboard,classifying and routing incidents by severity and tower while owning accurate, complete documentation of every action. Success is measured byfirst-response SLA adherence, triage accuracy, and runbook resolution rate.The role runs on arotational scheduleto sustain 24×7 coverage.

Key Responsibilities

Alert Monitoring & First Response

  • Monitor systems-estate health signals (compute, storage, OS, virtualization) from a consolidated single-pane-of-glass dashboard

  • Acknowledge alerts and incidents within first-response SLA targets

  • Perform initial review – read the alert, check recent changes, and review the CMDB before acting

Triage & Severity Classification

  • Classify each incident by severity and customer impact, confirming the owning tower (Systems / Network / Security / cross-tower)

  • Route and assign tickets accurately to the correct tower or escalation path

  • Sustain triage accuracy at or above the GIOC go-live standard

  • Reassess severity as impact evolves, and declare higher when in doubt

Runbook-Driven Resolution

  • Match alert signatures to documented runbooks and execute permitted L2 actions

  • Resolve common, well-understood systems issues end to end within L2 scope

  • Document every step taken, with supporting evidence, in the incident ticket

Escalation & Handoff

  • Escalate unresolved or out-of-scope issues to L2/L3 with a clean, complete handoff (symptom · evidence collected · steps tried · current state)

  • Flag missing or ineffective runbooks as runbook gaps for review and continuous improvement

  • Initiate and support major-incident (Sev-0 / Sev-1) bridges per the escalation matrix

  • Track escalations to closure and confirm service is restored before resolving the ticket

Operational Documentation & Communication

  • Maintain accurate shift logs, ticket updates, and incident timelines

  • Provide clear, timely status communication to stakeholders and customers per severity

  • Produce thorough shift-handover notes for seamless continuity

Continuous Improvement

  • Identify recurring alerts and noise as candidates for tuning and automation

  • Contribute to and improve runbooks and knowledge-base articles

  • Participate in post-incident reviews (PIR) and trend reviews

Qualifications & Experience

  • Graduate/Engineerin a relevant field (B.E./B.Tech, or equivalent)

  • 5-8 years of experiencein systems administration, NOC/SOC, or IT operations, includinghands-on Linux and/or Windows server administration

  • Working knowledge of operating systems – processes, services, and log analysis (Linux and/or Windows)

  • Familiarity with compute, storage, and virtualization concepts – VMs, hypervisors, and basic networking

  • Understanding of incident-management fundamentals and severity-based triage

  • Exposure to monitoring/observability tools and ticketing systems (alerting consoles, ITSM)

  • Strong written communication for accurate documentation and clean escalation handoffs; willingness and ability to work a rotational 24×7 shift model, including nights, weekends, and holidays

  • Proficient in English verbal and written communication

Inclusion & Privacy

We are an equal opportunity employer and are committed to creating an inclusive environment for all employees. We welcome applications from individuals of all backgrounds and experiences, and we prohibit discrimination based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected status under applicable laws. Vultr will consider qualified applicants with arrest or conviction records in accordance with applicable laws and will not conduct a background check until after an offer of employment has been extended and accepted.

We also take your privacy seriously. We handle personal information responsibly and follow applicable laws, including U.S. privacy rules and India’s Digital Personal Data Protection Act, 2023. Your data is used only for legitimate business purposes and is protected with proper security measures.

Where allowed by law, applicants may request details about the data we collect, access or delete their information, withdraw consent for its use, and opt out of nonessential communications. For more details, please see ourPrivacy Policy.

Apply Now

You'll be redirected to the company's application page

Requirements

  • ITIL V4 Foundation certification or equivalent ITSM knowledge
  • Linux certification (RHCSA / LFCS) or a cloud-fundamentals certification
  • Scripting familiarity (Bash, Python, or PowerShell) for routine automation
  • Exposure to cloud platforms and GPU / high-density infrastructure
  • Prior experience in a 24×7 NOC, SOC, or command-center environment, and tools such as JIRA, Confluence, PagerDuty, and observability platforms