
Job Description
Job Description
Site Reliability Engineer (SRE II) | 100% Onsite | Berkeley, CA | $80/hr | 1-Year Contract
Important Notes:
- Permanent overnight schedule of midnight to 8:00 a.m., five days per week
- 100% onsite in Berkeley, California
- U.S. Citizens only
- No third-party agencies, Corp-to-Corp (C2C), or subcontracting arrangements
We are seeking a Site Reliability Engineer (SRE II) to support Lawrence Berkeley National Laboratory‘s National Energy Research Scientific Computing Center (NERSC). As part of a 24×7 operations team, you‘ll help maintain the reliability and performance of critical high-performance computing infrastructure that supports scientific research and discovery. This role is ideal for an operations-focused engineer with strong Linux troubleshooting skills and experience in monitoring, alerting, incident response, automation, and production support.
What You‘ll Do:
- Monitor computing, storage, network, and facility systems
- Review and respond to production alerts and incidents
- Troubleshoot issues across Linux, applications, storage, networking, and infrastructure
- Develop and maintain automation, monitoring, and alerting solutions
- Build tools and integrations that support operational workflows
- Support incident management and documentation through ServiceNow
- Collaborate across technical teams to improve reliability and operational processes
- Perform periodic data center walkthroughs to verify environmental, cooling, and power systems
Qualifications:
Required
- Experience supporting Linux-based production systems in a 24×7 operations, SRE, NOC, data center, or similar environment
- Strong Linux administration and command-line experience
- Experience troubleshooting production issues from alert through resolution
- Experience developing tools or automation using Python, Perl, Java, C, C++, Bash, or similar
- Experience with monitoring, alerting, and operational support workflows
Preferred
- Experience in Site Reliability Engineering (SRE), DevOps, NOC, Systems Administration, Infrastructure Operations, or Platform Engineering
- ServiceNow experience
- Experience with Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, or similar monitoring platforms
- Familiarity with IT Service Management (ITSM) best practices
- Experience supporting HPC, research computing, scientific computing, or other mission-critical environments
- Experience developing automation tools; AI-driven automation experience is a plus
#zr