Lead Infrastructure Engineer, Managed Kubernetes Platform (MEKS)
Get to Know the Role
The Lead Infrastructure Engineer – Kubernetes & Service Mesh is a lead technical authority within the MEKS platform team. This role is responsible for designing, building, and operating the high-scale container and service mesh infrastructure. This infrastructure powers critical backend services across Grab. Reporting directly to the Platform Engineering Manager, this hands-on lead role focuses entirely on deep technical execution, platform architecture, system reliability, and developer experience. You will serve as the primary architect and technical mentor for the platform, driving multi-cluster AWS EKS strategies, Istio service mesh implementations, and cross-team developer enablement. This is an onsite position based in the Jakarta office.
The Critical Tasks You Will Perform Architecture & Platform Engineering
- Service Mesh Implementation involves designing and deploying production-grade Istio service mesh infrastructure. This infrastructure consists of a data plane and control plane. It requires managing various aspects, including mTLS, traffic management, canary/blue-green deployments, rate limiting, and circuit breaking, to optimize the service mesh.
- Architect and manage large-scale AWS EKS clusters, driving multi-cluster/multi-region topologies, node-pool optimization, custom scheduler strategies, and automated cluster lifecycles.
- Establish standardized cluster templates, namespaces, network policies, RBAC, and policy-as-code guardrails across the organization.
Developer Experience & Self-Service Tooling
- Build and maintain self-service APIs, GitOps workflows, and automation that simplify onboarding, deployments, and traffic management for internal product engineering teams.
- Partner directly with application teams to guide architectural choices, streamline workload migrations to MEKS, and accelerate time-to-market.
- Create comprehensive platform documentation, golden path templates, runbooks, and reference architectures to foster engineering self-sufficiency.
Reliability, Observability & Incident Response
- Implement telemetry, metrics, distributed tracing (OpenTelemetry/Jaeger), and centralized logging across clusters, Envoy proxies, and application workloads.
- Establish and uphold technical SLOs/SLAs, driving proactive capacity management, resource right-sizing, and performance tuning for peak availability and optimal infrastructure efficiency.
- Act as a senior technical point of escalation for complex platform outages, lead root-cause analyses , and execute corrective actions to prevent recurrence.
Technical Mentorship & Collaboration
- Mentor mid-level and senior engineers on the team in modern Kubernetes, mesh architecture, and distributed systems best practices.
- Collaborate with Security, SRE, Networking, and Cloud Infrastructure teams to ensure platform security, regulatory compliance, and seamless integration.
What Essential Skills You Will Need
- 10+ years of professional experience in infrastructure engineering, DevOps, or cloud platforms, combining a proven track record of designing, deploying, and managing production Kubernetes environments with demonstrated success in mentoring and leading technical teams.
- Expert-level knowledge of Kubernetes architecture and container technologies, including experience with configuration managers (Helm, Kustomize), service meshes (Istio, Linkerd), and cloud-native deployment patterns.
- Hands-on experience across major cloud platforms (AWS, GCP, or Azure) and multi-cloud management, leveraging Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation.
- Proven ability to design and implement automated CI/CD pipelines, deployment automation, and modern GitOps practices.
- Solid foundation in Linux/Unix system administration and networking principles, paired with proficiency in scripting languages like Python, Bash, or Go.
- Practical experience with monitoring and observability platforms (Prometheus, Grafana, ELK Stack) and a strong understanding of infrastructure security and compliance frameworks.
About Grab and Our Workplace
Grab is Southeast Asia's leading superapp. From getting your favourite meals delivered to helping you manage your finances and getting around town hassle-free, we've got your back with everything. In Grab, purpose gives us joy and habits build excellence, while harnessing the power of Technology and AI to deliver the mission of driving Southeast Asia forward by economically empowering everyone, with heart, hunger, honour, and humility.
Life at Grab
We care about your well-being at Grab, here are some of the global benefits we offer:
- We have your back with Term Life Insurance and comprehensive Medical Insurance.
- With GrabFlex, create a benefits package that suits your needs and aspirations.
- Celebrate moments that matter in life with loved ones through Parental and Birthday leave, and give back to your communities through Love-all-Serve-all (LASA) volunteering leave
- We have a confidential Grabber Assistance Programme to guide and uplift you and your loved ones through life's challenges.
- Balancing personal commitments and life's demands are made easier with our FlexWork arrangements such as differentiated hours
What We Stand For At Grab
We are committed to building an inclusive and equitable workplace that provides equal opportunity for Grabbers to grow and perform at their best. We consider all candidates fairly and equally regardless of nationality, ethnicity, race, religion, age, gender, family commitments, physical and mental impairments or disabilities, and other attributes that make them unique.
Meet the Engineering team
From software developers to security specialists, our people thrive on experimenting with fresh ideas and implementing them for users. Solve real-world problems, find purpose in your work and work across borders to help drive Southeast Asia forward.




