Software Engineering Group Manager – Site Reliability
Jobtailor
Alabama, United States Full Time Technology Jobs United States New
Job Description
Responsibilities
- Lead during moments that matter; own major incident response for high-impact (P1/P2) events, ensuring rapid resolution and clear communication.
- Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across: applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations; ensure efficient log, metric, and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering, infrastructure, and vendor teams.
- Drive root cause and real fixes; champion a culture of accountability through deep root cause analysis; eliminate repeat issues by driving permanent, systemic solutions and turn data and trends into actionable improvements.
- Shape reliability at scale; Define and evolve reliability strategy across availability, resiliency, and performance; lead improvements in uptime, MTTR, and overall system health and partner with engineering to embed reliability into system design.
- Modernize operations; advance observability with best-in-class monitoring, alerting, and event management; leverage tools like Dynatrace, BigPanda, and Logscale to enable proactive detection and drive automation to reduce manual effort and create self-healing systems.
- Ensure safe and reliable change; oversee change and release governance to enable fast and safe deployments; improve change success rates, reduce production defects, and lead post-release reviews that fuel continuous improvement.
- Lead a Global 24x7 Operation; manage distributed teams supporting critical systems around the clock; create seamless handoffs and strong operational discipline across regions and elevate team performance, engagement, and growth. Build a trusted, compliant environment; ensure alignment with enterprise governance, audit, and regulatory standards and strengthen risk management, controls, and operational documentation.
Requirements
- 8+ years of related experience and 5+ years of management experience.
- Proven leadership experience in Production Operations, SRE, or Infrastructure Engineering.
- Deep expertise in incident, problem, and change management within complex environments.
- Passion for building reliable, scalable, customer-centric platforms.
- Track record of improving operational metrics and leading high-performing teams.
- Strong executive presence and communication skills.
- Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.
Posted August 3, 2026