Vultr Cares
Medical Insurance stipend paid annually
9 Company-Paid Holidays
Generous Leave Policy + 1 month paid sabbatical every 5 years + Anniversary Bonus each year
Professional Development Reimbursement
Internet reimbursement
Fitness membership reimbursement
Company paid Wellable subscription
Join Vultr
Vultr is expanding its India presence and is building its firstGlobal Integrated Operations Command Center (GIOC) in Chennai– the 24×7 nerve center for monitoring, triage, and first response across Vultr’s globalcloud, systems, network, and security operations.
We are seekingSystems Engineersto serve as the first line of response for Vultr’s systems estate –compute, storage, operating systems, and virtualization.This is a frontline operations role for engineers who thrive in a fast-paced, alert-driven environment and want to grow into deeper SRE, systems, and platform engineering careers. You must be comfortable working on a rotational shift model (24×7) – including nights, weekends, and holidays, to ensure we have full coverage.
Role Overview
TheSystems Engineeris the entry point of the GIOC incident lifecycle. You monitor health signals across Vultr’s systems estate, acknowledge and triage alerts withindefined first-response SLAs,execute documented runbooks to resolve common issues, and escalate cleanly to Senior Systems Engineer when an issue falls outside L1 scope. The role ensuresfast, accurate first responsethat protects customer experience and keeps incidents moving toward resolution.
You operate from aconsolidated single-pane-of-glass dashboard, classifying and routing incidents by severity and tower while owning accurate, complete documentation of every action. Success is measured byfirst-response SLA adherence, triage accuracy, and runbook resolution rate.The role runs on arotational scheduleto sustain 24×7 coverage.
Key Responsibilities
Alert Monitoring & First Response
Monitor systems-estate health signals (compute, storage, OS, virtualization) from a consolidated single-pane-of-glass dashboard
Acknowledge alerts and incidents within first-response SLA targets
Perform initial review – read the alert, check recent changes, and review the CMDB before acting
Triage & Severity Classification
Classify each incident by severity and customer impact, confirming the owning tower (Systems / Network / Security / cross-tower)
Route and assign tickets accurately to the correct tower or escalation path
Sustain triage accuracy at or above the GIOC go-live standard
Reassess severity as impact evolves, and declare higher when in doubt
Runbook-Driven Resolution
Match alert signatures to documented runbooks and execute permitted L1 actions
Resolve common, well-understood systems issues end to end within L1 scope
Document every step taken, with supporting evidence, in the incident ticket
Escalation & Handoff
Escalate unresolved or out-of-scope issues to Senior Engineers with a clean, complete handoff (symptom · evidence collected · steps tried · current state)
Flag missing or ineffective runbooks as runbook gaps for review and continuous improvement
Initiate and support major-incident (Sev-0 / Sev-1) bridges per the escalation matrix
Track escalations to closure and confirm service is restored before resolving the ticket
Operational Documentation & Communication
Maintain accurate shift logs, ticket updates, and incident timelines
Provide clear, timely status communication to stakeholders and customers per severity
Produce thorough shift-handover notes for seamless continuity
Continuous Improvement
Identify recurring alerts and noise as candidates for tuning and automation
Contribute to and improve runbooks and knowledge-base articles
Participate in post-incident reviews (PIR) and trend reviews
Qualifications & Experience
Graduate/Engineerin a relevant field (B.E./B.Tech, or equivalent)
3-5 years of experiencein systems administration, preferably in NOC/SOC, or IT operations
Good understanding of Linux/Windows
Working knowledge of operating systems – processes, services, and log analysis (Linux/Windows)
Familiarity with compute, storage, and virtualization concepts – VMs, Hypervisors, and basic networking
Understanding of incident-management fundamentals and severity-based triage
Exposure to monitoring/observability tools and ticketing systems (alerting consoles, ITSM)
Strong written communication for accurate documentation and clean escalation handoffs; willingness and ability to work a rotational 24×7 shift model, including nights, weekends, and holidays
Proficient in English verbal and written communication
Inclusion & Privacy
We are an equal opportunity employer and are committed to creating an inclusive environment for all employees. We welcome applications from individuals of all backgrounds and experiences, and we prohibit discrimination based on race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected status under applicable laws. Vultr will consider qualified applicants with arrest or conviction records in accordance with applicable laws and will not conduct a background check until after an offer of employment has been extended and accepted.
We also take your privacy seriously. We handle personal information responsibly and follow applicable laws, including U.S. privacy rules and India’s Digital Personal Data Protection Act, 2023. Your data is used only for legitimate business purposes and is protected with proper security measures.
Where allowed by law, applicants may request details about the data we collect, access or delete their information, withdraw consent for its use, and opt out of nonessential communications. For more details, please see ourPrivacy Policy.