Site Reliability Engineer (SRE II)

  • Full Time
  • Berkeley
  • 0.000000 - 0.000000
Perfect Timing Personnel Services, Inc.

Job Description

Job Description

Site Reliability Engineer (SRE II) | 100% Onsite | Berkeley, CA | $80/hr | 1-Year Contract

Important Notes:
 

  • Permanent overnight schedule of midnight to 8:00 a.m., five days per week
  • 100% onsite in Berkeley, California
  • U.S. Citizens only
  • No third-party agencies, Corp-to-Corp (C2C), or subcontracting arrangements

We are seeking a Site Reliability Engineer (SRE II) to support Lawrence Berkeley National Laboratory‘s National Energy Research Scientific Computing Center (NERSC). As part of a 24×7 operations team, you‘ll help maintain the reliability and performance of critical high-performance computing infrastructure that supports scientific research and discovery. This role is ideal for an operations-focused engineer with strong Linux troubleshooting skills and experience in monitoring, alerting, incident response, automation, and production support.



What You‘ll Do:

  • Monitor computing, storage, network, and facility systems
  • Review and respond to production alerts and incidents
  • Troubleshoot issues across Linux, applications, storage, networking, and infrastructure
  • Develop and maintain automation, monitoring, and alerting solutions
  • Build tools and integrations that support operational workflows
  • Support incident management and documentation through ServiceNow
  • Collaborate across technical teams to improve reliability and operational processes
  • Perform periodic data center walkthroughs to verify environmental, cooling, and power systems

Qualifications:
Required

  • Experience supporting Linux-based production systems in a 24×7 operations, SRE, NOC, data center, or similar environment
  • Strong Linux administration and command-line experience
  • Experience troubleshooting production issues from alert through resolution
  • Experience developing tools or automation using Python, Perl, Java, C, C++, Bash, or similar
  • Experience with monitoring, alerting, and operational support workflows


Preferred

  • Experience in Site Reliability Engineering (SRE), DevOps, NOC, Systems Administration, Infrastructure Operations, or Platform Engineering
  • ServiceNow experience
  • Experience with Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, or similar monitoring platforms
  • Familiarity with IT Service Management (ITSM) best practices
  • Experience supporting HPC, research computing, scientific computing, or other mission-critical environments
  • Experience developing automation tools; AI-driven automation experience is a plus


#zr