- Posted On: 29/09/2026
Site Reliability Engineer
- Makati City | Work from home
- Site Reliability Engineer
- Full-Time
Client: AI-led property operations platform for multifamily real estate operators
A Site Reliability Engineer applies software engineering to operations, owning the reliability, performance and resilience of our production systems as a first-class engineering discipline. They define and defend the service levels our users depend on, reduce operational toil through automation, and lead how we respond to and learn from incidents — so we can keep shipping delightful software quickly and safely
Strategic Responsibilities
- Define and evolve the reliability strategy — service level indicators (SLIs), objectives (SLOs) and error budgets — partnering with product and engineering to balance feature velocity against reliability.
- Optimise reliability by managing scaling, capacity metrics, disaster recovery, failure and recovery processes.
- Embed observability into the development flow and tooling, ensuring automatic coverage of new applications and service layers
- Provide coaching, mentoring and reliability expertise across multiple teams and stakeholders.
Daily Responsibilities
- Own and operate the product infrastructure for availability, latency, scalability and cost, improving resilience and safe, repeatable deployment.
- Build and maintain observability — metrics, logging, tracing and SLO-based alerting — so issues are detected and understood before they reach users.
- Participate in the on-call rotation: respond to, mitigate and resolve production incidents, and lead or contribute to postmortems.
- Automate repetitive operational and deployment tasks to eliminate toil and reduce manual error.
- Partner with development teams on production-readiness reviews, capacity planning and safe rollout practices, holding the line on reliability standards consistently and in line with organisation standards.
- Fully participate in and provide feedback on the company’s engineering and operational practices, driving improvements as assigned.
- Respond to incidents while owning the infrastructure and reliability domain, and proactively lead restoration of infrastructure and scaling failures.
- Build and refine reporting and monitoring dashboards to manage critical application performance and availability.
Capabilities
Reliability Engineering
- Define, measure and defend SLIs, SLOs and error budgets, and use them to inform prioritisation and release decisions.
- Perform capacity planning and load/failure analysis to ensure systems scale predictably and degrade gracefully.
Infrastructure and Architecture
- Provide guidance on the technical feasibility, resilience and scope of engineering and re-architecting needed to meet reliability and performance goals.
- Automate the repetitive tasks required to maintain a secure, consistent and up-to-date operational environment.
- Contribute to securing the product and infrastructure so the data and operations of our users and the organisation are protected from harm, both malicious and inadvertent.
Observability and Monitoring
- Instrument systems and build dashboards and alerting that reflect real user experience, not just machine health.
- Ensure alerts are actionable and tied to SLOs, minimising noise and alert fatigue.
Incident Response and Operational Excellence
- Actively troubleshoot and lead the response to production issues raised in the course of serving our users.
- Run and contribute to blameless postmortems, maintain runbooks, and ensure follow-up actions are tracked to completion.
Automation and Delivery
- Be an enabler for software and operational engineers, integrating modern tools and practices into our environment, including:
- Infrastructure as code and configuration management.
- Containerisation and cloud-native techniques.
- Optimal, safe CI/CD environments and progressive delivery methodologies.
Technology
- Collaborate across the engineering team to ensure software delivered meets the highest reliability standards.
- Make technology recommendations that justifiably meet business and reliability goals.
- Maintain an in-depth understanding of technologies and stay abreast of current industry trends and emerging practices.
Process
- Promote and drive the adoption of SRE standards, guidelines and best practices.
- Review, propose and implement improvements to reliability and operational practices, including the error budget policy.
- Implement and improve the operational and tooling toolset.
Coaching / Mentoring
- Share broad knowledge of reliability practices, technologies and architectures, and function as an SRE mentor across the product engineering teams.
- Mentor other engineers toward consistent, robust and resilient outcomes.
Skills / Knowledge / Experience / Qualifications
- Relevant tertiary qualifications (e.g. Bachelor / Grad Dip. Science / Computer Science / Software Engineering, IT certification or similar) or equivalent practical experience.
- Either
- Demonstrated ability and broad experience in software development, with a strong interest in operations and reliability, or
- Demonstrated ability and broad experience in operations, with proven software development experience.
- Understanding of SRE principles — SLIs/SLOs, error budgets, toil reduction and blameless incident management.
- Experience operating production systems and participating in an on-call rotation.
- Preferably knowledgeable in or familiar with:
- Linux
- GCP, Kubernetes and containerisation
- PostgreSQL (or any relational SQL)
- Infrastructure as code (e.g. Terraform) and CI/CD tooling
- Observability tooling (e.g. Prometheus, Grafana, OpenTelemetry, or similar)
- A systems/automation language (e.g. Go or Python)
Personal Attributes / Qualities
- Positive work attitude.
- Clear spoken and written communication, able to interact professionally with a diverse group of clients and staff.
- Calm, methodical and decisive under pressure, particularly during incidents.
- Data-driven, with a blameless, systems-thinking approach to failure.
- Delivers high-quality technical outcomes in line with estimates, meeting product and customer needs.
- A high standard in engineering and operational practices.
- Ability to use initiative, and to handle and prioritise multiple requests from different sources.
- Be
- able to work both independently and within a team
- able to solve problems completely and quickly
- able to quickly learn a subject matter area
- a constructively critical thinker
- Reliable
Work Details
- Shift: Monday to Friday: Between 6:00am- 3:00pm or 7:00am- 4:00pm PH Time; depending on business needs
- Location: *Work from Home Until Further Notice
- Status: Full-time Employment