About the role
Sigil runs in five regions today and inside dozens of customer VPCs. You’ll own the platform both run on: clusters, networking, deploys and the tooling that keeps them boring.
What you’ll do
Operate and evolve our Kubernetes clusters across AWS and GCP
Build the Helm charts and Terraform modules customers use to self-host Sigil
Automate region launches, failover drills and capacity planning
Improve deploy safety with canaries, rollbacks and good defaults
What we’re looking for
5+ years running Kubernetes in production at scale
Strong Terraform, networking and Linux fundamentals
A habit of turning incidents into automation
Calm on-call judgment and clear written postmortems
Nice to have
Experience with GPU scheduling or model-serving infrastructure
SOC 2 or ISO 27001 audit experience