You join the Platform team and own the reliability of the systems our AI agents run on. This is a reliability engineering role: you treat operations as a software problem, so where others run a manual procedure, you write the automation that makes it unnecessary. You define what "healthy" means in numbers, measure it, and hold the line on it in production. When something breaks, you bring the system back and then make sure it cannot break the same way twice.
Reliability & Operations
Define and own SLOs and SLIs for the platform and manage error budgets against them
Carry on-call, act as incident commander, and run blameless post-incident reviews that produce real follow-up
Run production readiness and capacity planning ahead of demand, not after the page fires
Platform & Infrastructure
Run and harden the Kubernetes platform (Helm, GitOps, service mesh) and the cloud underneath it (Terraform, multi-region)
Own observability: metrics, logs, and distributed tracing, so problems surface before users feel them
Automation & Efficiency
Eliminate toil through automation, self-healing systems, and automated remediation
Drive cost visibility and FinOps practice across cloud and LLM spend
Security
Bake security into the platform: least-privilege access, secrets management, policy as code, and vulnerability management
Expected AI Skills
At Blockbrain, we don't just talk about AI — we use it every day. In this role, you will:
Use coding agents to build automation, write infrastructure code, and reason through failure modes faster
Automate operational toil and incident workflows with AI in the loop
Collaborate with the product team to give real-world feedback on Blockbrain's own tools from an operator's perspective
Stay curious about emerging AI capabilities and apply them to platform and reliability work
We're not looking for AI experts — we're looking for people who are genuinely open to working with AI as a daily co-pilot.
