Senior LLMOps Engineer
PatSnap
We are looking for a Senior LLMOps Engineer with strong software engineering and infrastructure skills to join PatSnap’s Singapore team.
You will lead engineering initiatives for model training, inference services, and training data management, providing reliable infrastructure, supporting platforms, and technical support to AI/ML and product engineering teams. This is a hands-on individual contributor role with technical leadership responsibilities, contributing directly to the company’s AI transformation.
Key Responsibilities
- Lead the development of model training platforms and workflows, supporting resource scheduling, job management, and troubleshooting for training and fine-tuning workloads.
- Build and operate model inference capabilities, covering open-source model deployment and external LLM API integration, with centralized access, permission, quota, and request management.
- Build training data management capabilities, including dataset storage, versioning, quality checks, access control, and traceability.
- Develop supporting platforms and automation for lifecycle management, deployment, and monitoring of models, training jobs, and inference services.
- Improve GPU and compute resource utilization, and optimize the reliability, performance, and cost of training and inference workloads.
- Participate in a team on-call rotation providing 24/7 coverage for supported production services, resolve incidents, and drive continuous improvement.
- Collaborate with AI/ML, engineering, and security teams to provide platform support and technical guidance, enabling AI adoption across the company.
- At least 5 years of experience in software development, platform engineering, SRE, or MLOps, with hands-on experience in model training or inference infrastructure.
- Strong Python programming skills, with the ability to independently develop and maintain platform services and automation tools.
- Familiarity with engineering workflows for model training, fine-tuning, and inference, with experience deploying open-source models and integrating external LLM APIs.
- Familiarity with Linux, Kubernetes, GPU resource management, CI/CD, and observability technologies.
- Understanding of storage, versioning, and access control for training data and model artifacts.
- Proficiency in using AI tools to improve engineering efficiency, with strong technical judgment, problem-solving, and cross-team collaboration skills.
- Fluency in English; Chinese language proficiency is an advantage.
Preferred Qualifications
- Contributions to open-source projects related to AI infrastructure, model training, or inference tooling.
- Experience with distributed training, inference performance optimization, or GPU cluster management.
Skills
- Kubernetes
- Linux
- Llm
- Problem solving
- Python
- Site reliability
- Software engineering
- Teamwork


