Routiqo

Operations and Platform Engineer - Traffic Infrastructure

ByteDance

  • Singapore
  • Full-time

About the Team The Traffic Infrastructure team leverages unified platform capabilities to manage global edge infrastructure (China & Non-China), both self-built and third-party, providing standardized, compliant, scalable, and cost-effective traffic infrastructure capabilities for edge services. Our vision is to build a global edge traffic infrastructure platform and become the long-term cornerstone of ByteDance’s global edge business in terms of scale, performance, and cost.

Responsibilities

  • Responsible for the architecture design and engineering of the "network-traffic infrastructure" operation and maintenance & efficiency platform.
  • Responsible for the engineering of the CMDB, operation and maintenance automation, observability, stability, and change management systems for the "network-traffic infrastructure".
  • Responsible for the interactive design and system development of the efficiency tools for the "network-traffic infrastructure" business to improve the operational management efficiency.
  • Explore the application and implementation of intelligent operation and maintenance scenarios, and promote the intelligent evolution of system operation and maintenance.

Requirements

Minimum Qualification(s)

  • Bachelor's degree or above in computer science or a related field, with at least 3 years of relevant experience in R&D, system operation and maintenance, or SRE.
  • Solid foundation in computer theory, with proficiency in at least one programming language such as Go, C, Python, etc.
  • Possess strong analytical skills, excellent communication abilities, a strong sense of responsibility and team spirit.
  • Passionate about programming, with a strong thirst for knowledge, curiosity and ambition.

Preferred Qualification(s)

  • Have experience in system engineering of large-scale distributed systems, management platforms or operation and maintenance platforms.
  • Familiar with infrastructure architecture, and have a solid understanding of Kubernetes, edge computing, cloud networking, Load Balance, microservice architecture and other related technologies.
  • Have a solid understanding of distributed systems, microservices architecture, high availability, stability assurance, and emergency response systems.
  • Have certain exploratory experience in LLM large model and Agent development.

Skills

  • Analytical
  • Communication
  • Kubernetes
  • Llm
  • Networking
  • Python
  • Site reliability
  • Solution architecture