Routiqo

Principal Engineer (Platform Engineering)

PayNet

  • Malaysia
  • Full-time

Why PayNet / Why Now

  • Build the secure platform foundations behind PayNet’s fraud intelligence, data, microservices, and machine learning capabilities.
  • Shape a critical stage of the PayNet Secure Project as workloads, data volumes, and production demands continue to grow.
  • Work where platform decisions must balance speed, resilience, security, regulatory expectations, and long-term sustainability.
  • Join a role with the mandate to make sound technical trade-offs, challenge assumptions, and turn architecture into dependable production outcomes.

##

TL;DR

  • Own the secure AWS and Kubernetes platform that supports fraud, data, microservices, and machine learning workloads.
  • Lead hands-on decisions across Terraform, Atlantis, GitLab CI/CD, Helm, Python automation, observability, and release controls.
  • Drive reliability, scalability, security, auditability, and cost-conscious use of platform resources across Development, UAT, and Production.
  • Enable production-grade MLOps while reducing dependency on senior architects through clear judgment and accountable execution.

##

Why This Role Matters

  • Platform reliability directly affects the stability and responsiveness of systems supporting fraud intelligence and live model operations.
  • Secure, traceable infrastructure and deployment decisions are essential in a regulated and security-sensitive environment.
  • Growing data workloads require deliberate architecture that scales without adding unnecessary operational complexity.
  • Strong technical ownership will create faster decisions, clearer accountability, and more sustainable delivery across engineering and data teams.

##

What You Will Actually Do

  • Own and evolve secure, highly available AWS platforms across segregated Development, UAT, and Production environments.
  • Architect and operate Kubernetes clusters, deployment patterns, workload scaling, resource optimisation, containers, and Helm releases.
  • Build controlled infrastructure workflows with Terraform and Atlantis, preserving traceability of infrastructure and configuration changes.
  • Lead GitLab CI/CD design for automated build, test, security validation, promotion, deployment, rollback, and release governance.
  • Drive observability, incident response, root-cause analysis, performance improvement, resilience, and cost-effective resource use.
  • Partner with Data Science, ML, security, networking, and infrastructure teams to enable secure model deployment and lifecycle operations.

##

Examples of This Role in Practice

  • A deployment introduces instability in Production: lead diagnosis, decide the rollback path, and drive a sustainable fix rather than a temporary workaround.
  • A data workload grows beyond 100GB per week: evaluate scaling options, make the trade-off explicit, and evolve the platform without over-engineering it.
  • A new tool could accelerate delivery but open-source adoption is restricted: assess the security and governance implications and recommend a viable path.
  • A live model needs lower latency and stronger monitoring: align platform, ML, and observability decisions to improve production performance and control.
  • An existing architecture decision is unclear: review the standards and decision records, challenge assumptions constructively, and document the chosen direction.

##

What Will Help You Succeed

  • At least 8 years of relevant experience in DevOps, Cloud Engineering, Platform Engineering, Site Reliability Engineering, or related infrastructure roles.
  • Deep production experience with AWS, Kubernetes, Terraform, GitLab CI/CD, Helm, Python automation, and mission-critical systems.
  • Strong judgment in cloud architecture, networking, IAM, secrets management, environment segregation, deployment controls, and security trade-offs.
  • Practical strength in monitoring, alerting, troubleshooting, performance optimisation, incident response, and root-cause analysis.
  • Ability to explain why systems are designed as they are, question assumptions, decide with incomplete information, and communicate clearly with stakeholders.
  • Useful exposure includes regulated environments, large-scale or distributed data processing, and production MLOps tools or lifecycles.

APPLY

Skills

  • AWS
  • Devops
  • Gitlab
  • Gitlab ci
  • Helm
  • Kubernetes
  • Python
  • Site reliability
  • Terraform