Director, Live Operations
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Experience: 8+ years in production operations, cloud/platform operations, DevOps, SRE, operations engineering, data platform operations, or a comparable high-availability technical environment; 4+ years leading technical operations teams
Core Areas: 24x7 Production Operations, AWS, Incident & Problem Management, Terraform, DevOps Automation, Observability, Data Pipelines, SLA/SLO Management
Compensation: $180,000-$210,000 per year
About the Role
The Director, Live Operations leads the technical production-operations function responsible for the availability, reliability, supportability, and continuous operation of managed customer data services. This role owns the 24x7 operating model for production data workflows and serves as the accountable leader for major operational incidents, service restoration, problem management, and continuous reliability improvement.
The position leads an operations-engineering organization spanning cloud production support, data-pipeline operations, observability, infrastructure automation, workflow recovery, secure data movement, access and connectivity, operational tooling, and automation. This is a hands-on technical leadership role requiring the ability to understand and challenge engineering approaches across AWS, Terraform and infrastructure-as-code, data pipelines, APIs, job orchestration, monitoring, and production automation.
What You'll Do
24x7 Production Reliability & Service Ownership- Own the 24x7 operational health and reliability of managed production data services and workflows.
- Serve as the accountable leader and escalation owner for Sev-1 and Sev-2 production incidents.
- Define SLAs and SLOs, escalation paths, on-call and coverage models, service-health measures, and operational performance expectations.
- Drive service restoration, communication coordination, and corrective-action follow-through.
- Own incident and problem-management disciplines, including severity definitions, incident command, escalation, root cause analysis, post-incident reviews, and corrective actions.
- Track recurring failures and use MTTA, MTTR, availability, incident volume, and recurrence metrics to drive systemic reliability improvements.
- Provide technical leadership for production services operating in AWS and related enterprise environments.
- Partner with Engineering on infrastructure-as-code using Terraform or comparable technologies, including repeatable configuration, deployment, and recovery.
- Guide operational troubleshooting involving IAM, networking and connectivity, secure file transfer, storage, compute, logging, monitoring, and cloud dependencies.
- Lead an automation-first strategy to eliminate repetitive manual work, fragile handoffs, and key-person dependencies.
- Drive scripting, orchestration, automated validation, job recovery, exception handling, and self-healing patterns where appropriate.
- Partner with Data Engineering on CI/CD, APIs, ETL and data pipelines, file movement, deployment and support patterns, and production automation.
- Apply AI-assisted monitoring, troubleshooting, documentation, and workflow automation where appropriate.
- Establish monitoring, logging, alerting, and operational dashboards that provide actionable visibility into production health.
- Define production-readiness gates for workflows transitioning from implementation or engineering into Live Operations.
- Require current runbooks, SOPs, recovery procedures, escalation paths, dependency maps, ownership, and cross-trained coverage before production handoff.
- Define clear operating boundaries among Live Operations, Data Engineering, Data Management, Clinical Intelligence, Product, and IT.
- Coordinate technical response across teams during production incidents and complex operational issues.
- Build and lead a geographically distributed technical operations team with a culture of urgency, transparency, documentation, collaboration, and measurable improvement.
Qualifications
Required Experience
- 8+ years of experience in production operations, cloud/platform operations, DevOps, SRE, operations engineering, data platform operations, or a comparable high-availability technical environment.
- 4+ years of experience leading technical production operations, DevOps, SRE, platform-support, or operations-engineering teams.
- Demonstrated accountability for business-critical production systems operating under extended-hours or 24x7 support models.
- Demonstrated leadership of Sev-1 and Sev-2 or equivalent major incidents, including incident command, service restoration, root cause analysis, problem management, and corrective-action follow-through.
- Experience leading geographically distributed technical teams and establishing effective on-call, escalation, and coverage models.
Required Skills
- Strong working knowledge of AWS production environments, including IAM, networking and connectivity, logging and monitoring, cloud dependencies, and operational troubleshooting.
- Demonstrated experience with Terraform or comparable infrastructure-as-code technologies and repeatable infrastructure deployment and recovery practices.
- Strong technical understanding of APIs, ETL and data pipelines, secure file transfer, workflow and job orchestration, SQL, automation and scripting, and production integration patterns.
- Demonstrated ability to establish observability, monitoring, alerting, runbooks, operational dashboards, and measurable reliability practices.
- Demonstrated success automating manual production processes and reducing key-person dependencies through tooling, scripting, orchestration, or platform improvements.
- Experience managing availability, SLA/SLO attainment, MTTA, MTTR, incident volume, recurring failures, and automation coverage.
- Ability to lead across Engineering, Product, IT, implementation, and customer-facing organizations during production incidents and reliability initiatives.
- Technical credibility with engineers and the ability to engage deeply on architecture and production failure modes.
- Ability to remain structured and decisive during high-severity incidents while maintaining clear operational communication, documentation, and accountability.
- Metrics-driven approach focused on systemic reliability improvement and eliminating avoidable manual and key-person dependencies.
Preferred Qualifications
- Experience operating high-volume healthcare, financial services, SaaS, or other regulated production data environments.
- Experience with MuleSoft, Flatfile, or comparable enterprise integration and data-ingestion platforms.
- Experience with CI/CD tooling, cloud observability platforms, workflow orchestration, and automated recovery patterns.
- Experience applying AI-assisted tooling to production support, incident analysis, documentation, or operational automation.
- Knowledge of healthcare payer data, CMS submissions, EDI/X12, enrollment, claims, risk adjustment, or related healthcare domains.
Benefits
- Medical, dental, and vision benefits.
- 401(k) match.
- Generous paid time off plan.
Looking for more opportunities?
View All Jobs