Agent Algorithm Evaluation Engineer
Shopee
Team Description
Sea Group is establishing a brand-new, strategic AI department. This department is dedicated to exploring the transformative potential of generative AI in revolutionizing human connection, self-expression and communication diversity, and social interaction. We are building the next generation of AI-native applications and a comprehensive Model-as-a-Service (MaaS) product support system. Based on massive multi-country data, we are building a leading multilingual AI ecosystem from the ground up. We look forward to more outstanding talents joining us to build leading Southeast Asian multilingual models and explore innovative AI-native applications.
The AI application team focuses on the intersection of social connectivity and artificial intelligence. Our mission is to leverage LLMs to create digital personas that can act as personal assistants and social bridges. This team operates with a startup's agility backed by our Group's robust resources, aiming to define how humans interact in the AI era.
Job Description
- Evaluation Framework Design: Design and build multi-dimensional evaluation metric systems tailored to diverse business scenarios.
- Red-Team Testing & Safety Control: Simulate complex, extreme real-world user scenarios to conduct stress testing and red-teaming of AI models; identify and drive remediation of issues related to bias, hallucinations, policy-violating content, and values alignment.
- In-Depth Bad Case Analysis: Conduct root-cause analysis of erroneous model outputs and collaborate with algorithm engineers (Model/SFT/RLHF) to drive prompt optimization and fine-tuning improvements.
- Automated Evaluation Tooling: Leverage LLM-as-a-Judge methodologies to build automated evaluation pipelines that enhance product iteration efficiency.
- User Experience Insights: Conduct in-depth research into user psychology within social and companion-style products, translating subjective, qualitative experience into quantifiable metrics to continuously enhance product experience.
Requirements
- Proven experience enhancing automated evaluation workflows, including but not limited to conventional evaluation methods, LLM-as-a-Judge, and automated weakness mining.
- Proven experience enhancing automated data production pipelines, including but not limited to Self-Instruct, Evol-Instruct, and Synthetic Data RL.
- Strong software engineering fundamentals, with proficiency in Python and Shell scripting and hands-on experience in Linux environments.
- Master's degree or above in Computer Science or a related field, with at least 3 years' working experience.
Nice-to-Have
- Prior background in AI evaluation, QA automation, or NLP algorithm development at a major internet company.
- Extensive experience as a power user or developer of mainstream AI social products.
Skills
- Bash
- Linux
- Llm
- Nlp
- Python
- Quality assurance
- Software engineering

