Symbotic

Staff Senior Reliability Engineer

Remote
July 23, 2026
Apply Now
Deadline date:

Job Description

What we need

You’ll be the Senior Reliability Engineer who owns complex production investigations end to end — from root cause, through customer communication, through verified corrective action. You’ll be the face of our RCA process to customers and senior leadership, the analyst who turns incident trends into strategic improvements, and a mentor who raises the quality bar for the whole team.

What you’ll do

  • Lead high-impact RCA investigations for complex production incidents spanning software, infrastructure, industrial controls, and production SOP execution

  • Chair structured, blameless RCA reviews aligned with ITIL Problem Management — fact-based analysis, clear ownership, timely resolution

  • Serve as the customer-facing technical lead for RCA discussions, updates, and formal deliverables within defined SLA timelines

  • Analyze logs, telemetry, and incident trends across many issues to find recurring failure patterns — then influence teams across the organization to eliminate them

  • Present findings, risks, and recommendations to senior internal leadership and customer stakeholders, backed by data you own end to end

  • Drive Continuous Service Improvement initiatives that measurably reduce repeat incidents and investigation toil through automation, tooling, and reporting

  • Mentor teammates and elevate RCA quality standards as a senior individual contributor

What you’ll need

  • Minimum 8 years supporting complex, business-critical production environments, with a career centered on reliability

  • Minimum 5 years leading technical RCA, post-incident reviews, or ITIL-aligned Problem Management across software, infrastructure, systems, or industrial technology domains

  • Strong hands-on troubleshooting and data analysis across large-scale distributed systems, on-prem infrastructure, custom software, logs, telemetry, and incident datasets

  • A track record of regular, ongoing customer interaction — you can tell us who you worked with, at what level, and how often — and of earning trust with executives, technical and non-technical alike

  • Proven ability to run multiple high-priority investigations in parallel while influencing cross-functional teams, without direct authority, to close actions on time

  • Bachelor’s degree in a technical field, or equivalent practical experience

Bonus points

  • Hands-on experience with warehouse automation, robotics, industrial controls, or large-scale physical production environments (strongly preferred)

  • Kubernetes, VMware, Linux, Grafana, Prometheus, Zabbix, GitLab, and AI-assisted troubleshooting and analysis

  • Power BI, Tableau, reporting automation, or dashboard development for RCA analysis and leadership visibility

  • Depth in ITSM/ITIL practice — Incident Management, Problem Management, RCA, and Continuous Service Improvement

#LI-PA2

#LI-Remote