Site Reliability Engineer (SRE) – II

  • Full Time
  • Wisconsin
  • 0.000000 - 0.000000
Dormont Manufacturing Co



Description

Applicants must be currently authorized to work in the United States on a full‑time basis and are not eligible for visa sponsorship (F‑1, H‑1B, O‑1, TN, E‑3).

Are you a natural problem solver who thrives in high‑pressure situations and enjoys working across teams to keep systems running smoothly? We’re looking for a Site Reliability Engineer (SRE) Level II who brings technical expertise and strong communication to support and scale our critical systems.

This role is ideal for someone who:


  • Is personable and proactive, with a knack for jumping into complex issues and guiding troubleshooting conversations.
  • Can learn quickly, connect the dots across systems, and adapt to new technologies on the fly.
  • Enjoys working with others—whether it’s developers, support teams, or business stakeholders—to solve problems and improve reliability.

You’ll be part of a team that ensures our systems are resilient, scalable, and well‑supported. While technical depth is important, we especially want someone who can lead incident response, communicate clearly, and drive continuous improvement in systems and processes.

Key Responsibilities

Incident Response & Troubleshooting

  • Lead real‑time troubleshooting efforts for high‑impact production issues.
  • Collaborate across IT and engineering teams to resolve incidents quickly and effectively.
  • Provide mentorship and guidance to junior SREs and support staff.
  • Participate in on‑call rotations and act as an escalation point.

Automation & Infrastructure as Code

  • Build and maintain automation using tools like Terraform, Ansible, or CloudFormation.
  • Eliminate manual tasks and improve reliability through scripting and automation.

Monitoring & Observability

  • Build and optimize monitoring dashboards using tools like Prometheus, Dynatrace, Splunk etc.
  • Ensure visibility into system health and proactively detect issues.

Continuous Improvement

  • Drive improvements in deployment, monitoring, and incident response processes.
  • Champion best practices across the SRE and support teams.

Basic Qualifications

  • Bachelor’s degree in computer science, information technology, or related field.
  • 3+ years of experience in site reliability engineering, DevOps, systems administration, or related roles.

Preferred Qualifications

  • Strong troubleshooting and communication skills in production environments.
  • Experience supporting applications in both .NET and Spring Boot frameworks.
  • Familiarity with OpenShift, Windows Server, and hybrid deployment environments.
  • Proficiency in log analysis using SQL queries and Splunk.
  • Strong scripting skills (e.g., PowerShell, Bash, Python).
  • Familiarity with cloud platforms (AWS, GCP, etc.).
  • Hands‑on experience with observability tools (Dynatrace, Datadog, etc.).
  • Strong interpersonal skills and a customer‑focused mindset.

Exempt Status: Yes (not eligible for overtime pay).

Workplace Type: Office.

Our Approach to Office Workplace Type: Certain positions may be eligible for a flexible work arrangement, combining in‑office and work‑from‑home. Remote roles will also have the opportunity to come together in our offices for moments that matter. Specific work arrangements will be provided by the hiring team.



Huntington is an Equal Opportunity Employer.


#J-18808-Ljbffr