Site Reliability Engineer
  • Posted On: 29/09/2026

Site Reliability Engineer

  • Makati City | Work from home
  • Site Reliability Engineer
  • Full-Time
  • Apply Now

Client:  AI-led property operations platform for multifamily real estate operators

A Site Reliability Engineer applies software engineering to operations, owning the reliability, performance and resilience of our production systems as a first-class engineering discipline. They define and defend the service levels our users depend on, reduce operational toil through automation, and lead how we respond to and learn from incidents — so we can keep shipping delightful software quickly and safely

Strategic Responsibilities

  • Define and evolve the reliability strategy — service level indicators (SLIs), objectives (SLOs) and error budgets — partnering with product and engineering to balance feature velocity against reliability.
  • Optimise reliability by managing scaling, capacity metrics, disaster recovery, failure and recovery processes.
  • Embed observability into the development flow and tooling, ensuring automatic coverage of new applications and service layers
  • Provide coaching, mentoring and reliability expertise across multiple teams and stakeholders.

Daily Responsibilities

  • Own and operate the product infrastructure for availability, latency, scalability and cost, improving resilience and safe, repeatable deployment.
  • Build and maintain observability — metrics, logging, tracing and SLO-based alerting — so issues are detected and understood before they reach users.
  • Participate in the on-call rotation: respond to, mitigate and resolve production incidents, and lead or contribute to postmortems.
  • Automate repetitive operational and deployment tasks to eliminate toil and reduce manual error.
  • Partner with development teams on production-readiness reviews, capacity planning and safe rollout practices, holding the line on reliability standards consistently and in line with organisation standards.
  • Fully participate in and provide feedback on the company’s engineering and operational practices, driving improvements as assigned.
  • Respond to incidents while owning the infrastructure and reliability domain, and proactively lead restoration of infrastructure and scaling failures.
  • Build and refine reporting and monitoring dashboards to manage critical application performance and availability.

Capabilities

Reliability Engineering

  • Define, measure and defend SLIs, SLOs and error budgets, and use them to inform prioritisation and release decisions.
  • Perform capacity planning and load/failure analysis to ensure systems scale predictably and degrade gracefully.

Infrastructure and Architecture

  • Provide guidance on the technical feasibility, resilience and scope of engineering and re-architecting needed to meet reliability and performance goals.
  • Automate the repetitive tasks required to maintain a secure, consistent and up-to-date operational environment.
  • Contribute to securing the product and infrastructure so the data and operations of our users and the organisation are protected from harm, both malicious and inadvertent.

Observability and Monitoring

  • Instrument systems and build dashboards and alerting that reflect real user experience, not just machine health.
  • Ensure alerts are actionable and tied to SLOs, minimising noise and alert fatigue.

Incident Response and Operational Excellence

  • Actively troubleshoot and lead the response to production issues raised in the course of serving our users.
  • Run and contribute to blameless postmortems, maintain runbooks, and ensure follow-up actions are tracked to completion.

Automation and Delivery

  • Be an enabler for software and operational engineers, integrating modern tools and practices into our environment, including:
    • Infrastructure as code and configuration management.
    • Containerisation and cloud-native techniques.
    • Optimal, safe CI/CD environments and progressive delivery methodologies.

Technology

  • Collaborate across the engineering team to ensure software delivered meets the highest reliability standards.
  • Make technology recommendations that justifiably meet business and reliability goals.
  • Maintain an in-depth understanding of technologies and stay abreast of current industry trends and emerging practices.

Process

  • Promote and drive the adoption of SRE standards, guidelines and best practices.
  • Review, propose and implement improvements to reliability and operational practices, including the error budget policy.
  • Implement and improve the operational and tooling toolset.

Coaching / Mentoring

  • Share broad knowledge of reliability practices, technologies and architectures, and function as an SRE mentor across the product engineering teams.
  • Mentor other engineers toward consistent, robust and resilient outcomes.

Skills / Knowledge / Experience / Qualifications

  • Relevant tertiary qualifications (e.g. Bachelor / Grad Dip. Science / Computer Science / Software Engineering, IT certification or similar) or equivalent practical experience.
  • Either
    • Demonstrated ability and broad experience in software development, with a strong interest in operations and reliability, or
    • Demonstrated ability and broad experience in operations, with proven software development experience.
  • Understanding of SRE principles — SLIs/SLOs, error budgets, toil reduction and blameless incident management.
  • Experience operating production systems and participating in an on-call rotation.
  • Preferably knowledgeable in or familiar with:
    • Linux
    • GCP, Kubernetes and containerisation
    • PostgreSQL (or any relational SQL)
    • Infrastructure as code (e.g. Terraform) and CI/CD tooling
    • Observability tooling (e.g. Prometheus, Grafana, OpenTelemetry, or similar)
    • A systems/automation language (e.g. Go or Python)

Personal Attributes / Qualities

  • Positive work attitude.
  • Clear spoken and written communication, able to interact professionally with a diverse group of clients and staff.
  • Calm, methodical and decisive under pressure, particularly during incidents.
  • Data-driven, with a blameless, systems-thinking approach to failure.
  • Delivers high-quality technical outcomes in line with estimates, meeting product and customer needs.
  • A high standard in engineering and operational practices.
  • Ability to use initiative, and to handle and prioritise multiple requests from different sources.
  • Be
    • able to work both independently and within a team
    • able to solve problems completely and quickly
    • able to quickly learn a subject matter area
    • a constructively critical thinker
  • Reliable

 

Work Details

 

  • Shift: Monday to Friday: Between 6:00am- 3:00pm or 7:00am- 4:00pm PH Time; depending on business needs
  • Location: *Work from Home Until Further Notice
  • Status: Full-time Employment

Scroll to Top