Poslodavci će vas vidjeti u našoj bazi podataka i moći će sami ponuditi posao
  • Pretraživanje poslova
  • Omiljeno
  • Izradite životopis
    Novo
  • Upisi

Senior/Staff Kubernetes Infrastructure Engineer

Stalni radni odnos

fal

Responsibilities

  • Design, automate, validate, and deliver the complete lifecycle of customer compute environments from provisioning through upgrades, recovery, and decommissioning.
  • Use AI to automate and accelerate infrastructure delivery and operations.
  • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads.
  • Build and maintain Linux images and automated operating-system provisioning workflows.
  • Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
  • Design Kubernetes and data-center networking using Cilium, Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
  • Configure distributed and shared storage for high-performance workloads.
  • Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
  • Develop reusable tooling, standards, documentation, and runbooks.
  • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.

Requirements

  • At least 5 years of experience building and operating production Linux infrastructure.
  • Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, highly available control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
  • Experience with Linux virtualization using KVM/QEMU, libvirt, and VFIO device passthrough.
  • Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
  • Strong networking fundamentals covering TCP/IP, Layer 2/Layer 3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark.
  • Practical scripting experience and experience with configuration-management tools such as Ansible.
  • Ability to diagnose complex cross-layer infrastructure issues and drive technical decisions across teams.
  • Production Slurm experience is preferred.
  • Preferred experience includes NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX, hugepages, NUMA, CPU pinning, SR-IOV, and DPDK.
  • Preferred distributed-storage experience includes Ceph, Lustre, or Weka.
  • Preferred experience includes KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, and Nornir.
  • AI training, inference, or distributed GPU workload infrastructure experience is preferred.
  • Python or Go proficiency is preferred.
Oglas je objavljen prije 15 danaRok za prijavu traje do: 16. studenoga 2026.