Senior DevOps Engineer, Observability
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Employment Type: Full-Time
Experience: 5+ years in DevOps, SRE, or Platform Engineering; 2+ years at a B2B software startup
Core Areas: AWS, Kubernetes/EKS, Terraform, Helm, CI/CD Automation, Observability, Incident Response, Change Data Capture (CDC) and Event Streaming
Compensation: $155,000 – $185,000 annually, plus equity and bonus opportunities
About the Role
This opportunity is for a Senior DevOps Engineer, Observability to design, build, and operate reliable, scalable infrastructure supporting SaaS and on-premises software delivery. The role focuses on Infrastructure as Code, AWS cloud operations, CI/CD automation, observability, incident response, and internal platform tooling that enables engineering teams to deliver high-quality software efficiently and reliably.
The position combines hands-on infrastructure engineering with developer enablement, production reliability, and platform optimization. Responsibilities include improving monitoring, alerting, and service-level objectives (SLOs), strengthening security and SOC 2 compliance, and supporting systems involving Change Data Capture (CDC), event streaming, or large-scale multi-tenant observability. The role works closely with product engineering teams to improve infrastructure performance, cost efficiency, and developer self-service capabilities.
What You'll Do
- Design, build, and maintain reliable, scalable infrastructure systems supporting both SaaS and on-premises engineering environments.
- Operate and optimize AWS and other cloud infrastructure with a focus on security, performance, reliability, and cost efficiency.
- Develop and improve internal platform tooling, CI/CD automation using GitHub Actions, and developer self-service capabilities.
- Enhance observability and incident response systems through improved monitoring, alerting, and service-level objectives (SLOs).
- Collaborate with stream-aligned product engineering teams to understand their infrastructure requirements and continuously improve internal platforms.
- Implement, enforce, and improve infrastructure security and compliance standards, including SOC 2 controls.
- Create and maintain technical documentation, onboarding materials, and internal support processes.
- Participate in the on-call rotation to support production reliability and operational continuity.
Qualifications
Required Experience
- 5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering roles.
- 2+ years of experience working at a B2B software startup.
- Experience operating infrastructure in fast-paced environments with evolving priorities and ambiguous requirements.
Required Skills
- Strong hands-on experience with AWS infrastructure, including EC2, VPC, IAM, RDS, and related cloud services.
- Strong experience with Kubernetes and Amazon EKS for container orchestration and infrastructure operations.
- Proficiency with Infrastructure as Code technologies, including Terraform and Helm, for automated infrastructure provisioning and management.
- Experience designing, maintaining, and improving CI/CD pipelines and automation tooling, ideally using GitHub Actions.
- Familiarity with observability platforms such as Prometheus, Grafana, Mimir, Loki, or similar monitoring and logging technologies.
- Proficiency in Python, Go, or shell scripting for infrastructure automation and operational tooling.
- Experience with Change Data Capture (CDC) and event streaming systems, or experience scaling large, multi-tenant observability systems involving data ingestion, analysis, and alerting.
- Ability to work effectively in a fast-paced startup environment with changing priorities and technical ambiguity.
- Strong written and verbal communication skills, including the ability to produce clear technical documentation and collaborate with engineering teams.
Preferred Qualifications
- Experience working in SOC 2-compliant or security-focused environments.
- Experience with MQTT, AMQP, and other messaging technologies.
- Familiarity with network automation tooling and network infrastructure management ecosystems.
- Experience contributing to open-source software projects.
- Experience using AI-assisted development tools such as GitHub Copilot, ChatGPT, or Cursor.
Looking for more opportunities?
View All Jobs