Cloud Operations Engineer – Infrastructure - #1167254
TP-Link
Key Responsibilities
Design, build and maintain reliable, scalable and secure cloud-native infrastructure supporting large-scale production workloads.
Operate and optimise multi-account AWS environments, ensuring infrastructure is secure, repeatable and auditable.
Provision, administer and upgrade production Kubernetes clusters.
Manage Kubernetes networking, autoscaling, observability, capacity planning and day-to-day cluster operations.
Build and operate Kubernetes ecosystem components, including CRDs, Helm, HPA, Cluster Autoscaler, CoreDNS and Cluster API.
Develop and improve GitOps deployment workflows using FluxCD, ArgoCD or similar tools.
Manage and enhance Istio service mesh capabilities, including traffic routing, service discovery, resilience, security and service-to-service communication.
Automate infrastructure provisioning and configuration management using Terraform and related Infrastructure as Code tools.
Improve CI/CD pipelines, observability platforms and operational workflows using Go, Python or other appropriate technologies.
Establish and continuously improve reliability practices, including Service Level Objectives, Error Budgets, monitoring, alerting, incident response and post-incident reviews.
Investigate and resolve complex production issues involving AWS infrastructure, Kubernetes, Linux, networking and distributed services.
Collaborate with application engineering, architecture, security and platform teams to improve system reliability, scalability and operational efficiency.
Participate in a scheduled on-call rotation supporting production cloud infrastructure and Kubernetes platforms.
Requirements
Bachelor’s degree in Computer Science, Software Engineering, Information Technology or a related discipline.
At least two years of hands-on experience in cloud infrastructure, Kubernetes operations, platform engineering, Site Reliability Engineering, DevOps or a related area.
Strong knowledge of AWS, particularly EKS, IAM, VPC, EC2, S3 and associated networking and security capabilities.
Hands-on experience operating Kubernetes in a production environment, including:
Cluster provisioning and architecture
Workload orchestration
Networking and service discovery
Autoscaling and capacity management
Monitoring and troubleshooting
Familiarity with Kubernetes ecosystem tools such as CRDs, Helm, Cluster API, HPA, Cluster Autoscaler and CoreDNS.
Experience with GitOps tools such as FluxCD or ArgoCD.
Solid Linux administration and troubleshooting skills, including systemd, networking and performance analysis.
Experience with CI/CD pipelines and infrastructure automation using Terraform, Go, Python or similar tools.
Good understanding of reliability engineering practices, including SLOs, incident response, monitoring, alerting and post-incident reviews.
Strong analytical and problem-solving skills, with the ability to diagnose complex infrastructure issues across distributed systems.
Good communication and collaboration skills, with the ability to work effectively across engineering, security and architecture teams.
Willingness to participate in a scheduled on-call rotation.
Additional Advantages
Experience with NVIDIA device plugins, GPU scheduling or GPU workload operations in Kubernetes.
Experience with other public cloud platforms, particularly Azure or Alibaba Cloud.
Kubernetes certifications such as CKA, CKAD or CKS.
How to apply
To apply for this job you need to authorize on our website. If you don't have an account yet, please register.
Post a resumeSimilar jobs
iOS Mobile Developer (Senior)
Front Desk Customer Service cum Admin Officer (North/South Cluster)
General Insurance Admin (CGI / BCP / PGI)