Principal Engineer (Platform Engineering)
PayNet
Why PayNet / Why Now
- Build the secure platform foundations behind PayNet’s fraud intelligence, data, microservices, and machine learning capabilities.
- Shape a critical stage of the PayNet Secure Project as workloads, data volumes, and production demands continue to grow.
- Work where platform decisions must balance speed, resilience, security, regulatory expectations, and long-term sustainability.
- Join a role with the mandate to make sound technical trade-offs, challenge assumptions, and turn architecture into dependable production outcomes.
##
TL;DR
- Own the secure AWS and Kubernetes platform that supports fraud, data, microservices, and machine learning workloads.
- Lead hands-on decisions across Terraform, Atlantis, GitLab CI/CD, Helm, Python automation, observability, and release controls.
- Drive reliability, scalability, security, auditability, and cost-conscious use of platform resources across Development, UAT, and Production.
- Enable production-grade MLOps while reducing dependency on senior architects through clear judgment and accountable execution.
##
Why This Role Matters
- Platform reliability directly affects the stability and responsiveness of systems supporting fraud intelligence and live model operations.
- Secure, traceable infrastructure and deployment decisions are essential in a regulated and security-sensitive environment.
- Growing data workloads require deliberate architecture that scales without adding unnecessary operational complexity.
- Strong technical ownership will create faster decisions, clearer accountability, and more sustainable delivery across engineering and data teams.
##
What You Will Actually Do
- Own and evolve secure, highly available AWS platforms across segregated Development, UAT, and Production environments.
- Architect and operate Kubernetes clusters, deployment patterns, workload scaling, resource optimisation, containers, and Helm releases.
- Build controlled infrastructure workflows with Terraform and Atlantis, preserving traceability of infrastructure and configuration changes.
- Lead GitLab CI/CD design for automated build, test, security validation, promotion, deployment, rollback, and release governance.
- Drive observability, incident response, root-cause analysis, performance improvement, resilience, and cost-effective resource use.
- Partner with Data Science, ML, security, networking, and infrastructure teams to enable secure model deployment and lifecycle operations.
##
Examples of This Role in Practice
- A deployment introduces instability in Production: lead diagnosis, decide the rollback path, and drive a sustainable fix rather than a temporary workaround.
- A data workload grows beyond 100GB per week: evaluate scaling options, make the trade-off explicit, and evolve the platform without over-engineering it.
- A new tool could accelerate delivery but open-source adoption is restricted: assess the security and governance implications and recommend a viable path.
- A live model needs lower latency and stronger monitoring: align platform, ML, and observability decisions to improve production performance and control.
- An existing architecture decision is unclear: review the standards and decision records, challenge assumptions constructively, and document the chosen direction.
##
What Will Help You Succeed
- At least 8 years of relevant experience in DevOps, Cloud Engineering, Platform Engineering, Site Reliability Engineering, or related infrastructure roles.
- Deep production experience with AWS, Kubernetes, Terraform, GitLab CI/CD, Helm, Python automation, and mission-critical systems.
- Strong judgment in cloud architecture, networking, IAM, secrets management, environment segregation, deployment controls, and security trade-offs.
- Practical strength in monitoring, alerting, troubleshooting, performance optimisation, incident response, and root-cause analysis.
- Ability to explain why systems are designed as they are, question assumptions, decide with incomplete information, and communicate clearly with stakeholders.
- Useful exposure includes regulated environments, large-scale or distributed data processing, and production MLOps tools or lifecycles.
Skills
- AWS
- Devops
- Gitlab
- Gitlab ci
- Helm
- Kubernetes
- Python
- Site reliability
- Terraform

