Site Reliability Engineer: AI-Driven Kubernetes Automation

fal is seeking a senior platform engineer to own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads. You will build and maintain CI/CD pipelines, leverage AI to automate production analysis, and drive reliability improvements with automation, SLOs, and incident processes, collaborating across teams in Turkey.

Strong focus on observability with Prometheus, Grafana, Loki, and Datadog, and experience with

Back to blog