Responsibilities:
- Build, deploy, and operate backend systems that power AI-enabled features in production.
- Design and implement inference pipelines, orchestration layers, and service boundaries around AI models.
- Ensure production reliability through monitoring, logging, alerting, and incident response.
- Optimise latency and throughput across inference, caching, batching, and streaming workloads.
- Develop scalable, reliable backend systems capable of handling high-volume AI workloads with low latency and high throughput.
- Build stable and well-designed APIs that integrate seamlessly with frontend and machine learning systems.
- Monitor production environments, troubleshoot incidents, and implement continuous improvements to enhance system performance, scalability, and reliability.
Requirements:
- Proven experience in full stack software engineering, covering both frontend and backend development.
- Strong understanding of software architecture, system design, and API development.
- Experience working with Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), AI agents, or AI-powered applications.
- Ability to make sound engineering decisions in ambiguous and rapidly evolving environments.
- Strong ownership mindset with the ability to take features from concept through to production deployment.
- Comfortable working in a fast-paced environment with evolving business and technical requirements.
Technical Stack:
- Python
- Node.js
- PyTorch
- OpenAI, Anthropic, and open-source Large Language Models (LLMs)
- SQL and NoSQL databases
- Kubernetes
- Docker
To apply, please visit www.gmprecruit.com and search for Job Reference: 63R9X8Y5
To learn more about this opportunity, please contact Yingying at yingying.lai@gmprecruit.com
We regret that only shortlisted candidates will be notified.
GMP Technologies (S) Pte Ltd | EA Licence: 11C3793 | EA Personnel: Lai Yingying | Registration No: R1110239