Site Reliability Engineer | Hybrid

  • Full Time
  • Berkeley
  • 0.000000 - 0.000000
LTD GLOBAL, LLC

Job Description

Job Description

Hybrid — Berkeley, CA

Assignment: 10/26/2026 – 10/27/2027



$80/hr



Role Summary


 

As a Site Reliability Engineer on the Operations Technology team, you’ll be part of a round-the-clock crew keeping a national-scale HPC facility accessible, reliable, and secure. Working from advanced monitoring and data collection systems, you’ll proactively catch issues before they escalate, triage and resolve alerts across compute, storage, and network systems, and build the automation that makes the whole environment more resilient over time. You’ll also collaborate closely with cross-functional teams to coordinate maintenance, improve tooling, and ensure the infrastructure scales smoothly as demand grows, keeping the computational power behind critical scientific research running without interruption.


 

What You’ll Own

  • Monitor and triage alerts across computer, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution


 


What You Bring

  • Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA
  • Solid Linux/command-line (SSH) chops
  • Programming/scripting experience — Python, C, C++, Perl, or Java
  • A self-starter mindset — eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
  • Network security fundamentals (ACLs, firewalls)
  • Strong cross-team communication and collaboration skills


 


Nice to Have

  • Experience building or deploying Agentic AI / autonomous automation for technical workflows
  • ServiceNow implementation experience
  • ITSM best-practice know-how

nCompany Description


Great Organization!

Company Description

Great Organization!