Responsibilities:
- Manage and operate centralized monitoring and observability platforms across applications, databases, infrastructure (compute and storage), and network environments supporting 24/7 services.
- Continuously monitor system health using metrics, logs, dashboards, and alerts to proactively detect performance degradation, anomalies, and potential service issues.
- Perform initial incident triage, impact assessment, event correlation, and coordinate timely resolution with Application, Infrastructure, Network, Cloud, and Database teams.
- Design, implement, and continuously improve monitoring strategies, observability standards, alert thresholds, escalation policies, and monitoring frameworks.
- Develop and maintain real-time service health dashboards, operational reports, and trend analysis for management and key stakeholders.
- Support major incident management by providing technical diagnostics, operational visibility, and cross-functional coordination.
- Identify recurring incidents and reliability gaps, driving continuous improvements in monitoring coverage, system stability, and operational response.
- Monitor and optimise AWS cloud and infrastructure costs, including compute, storage, network connectivity, and data transfer expenses.
- Implement cost allocation, tagging strategies, budget monitoring, and alerting mechanisms to improve cloud cost governance.
- Prepare cost reports, dashboards, forecasts, and recommendations to optimise resource utilisation while balancing performance, reliability, and cost efficiency.
- Drive initiatives to enhance observability, automate monitoring and alerting workflows, and support the adoption of Site Reliability Engineering (SRE) practices, including SLIs, SLOs, and Error Budgets.
- Maintain accurate documentation covering monitoring architecture, dashboards, alerting rules, escalation procedures, and FinOps governance.
- Participate in major incident response, critical service monitoring, and after-hours support, including weekends and public holidays, when required.
Requirements:
- Degree or Diploma in Computer Science, Information Technology, Engineering, or a related discipline.
- Minimum 3 to 5 years of experience in IT Operations, Infrastructure Operations, System Monitoring, NOC, Service Assurance, or Cloud Infrastructure.
- Hands-on experience with enterprise monitoring and observability platforms.
- Experience supporting hybrid environments across on-premises infrastructure and AWS cloud.
- Good understanding of system, application, infrastructure, and network monitoring concepts.
- Experience with one or more monitoring platforms such as Amazon CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or equivalent solutions.
- Strong ability to analyse logs, metrics, alerts, and performance data to identify and resolve operational issues.
- Experience with AWS Cost Explorer, budgeting, tagging strategies, cloud cost optimisation, and FinOps best practices.
- Familiarity with ITIL processes, including Incident, Problem, Change Management, Service Level Management, and Observability principles.
- AWS Associate Certification or AWS FinOps Certified Practitioner certification will be an advantage.
- Strong analytical, troubleshooting, and problem-solving skills.
- Ability to correlate events across complex distributed systems and coordinate effectively with multiple technical teams.
- Strong communication, stakeholder management, and organisational skills.
- Self-motivated with a strong sense of ownership, accountability, and the ability to work independently.
- Willing to provide after-hours operational support when required.
To apply, please visit www.gmprecruit.com and search for Job Reference: X94RV4Y3
To learn more about this opportunity, please contact Yingying at yingying.lai@gmprecruit.com
We regret that only shortlisted candidates will be notified.
GMP Technologies (S) Pte Ltd | EA Licence: 11C3793 | EA Personnel: Lai Yingying | Registration No: R1110239