SRE Consultant – Site Reliability Engineering

Há 7 dias

São Paulo, State of São Paulo, Brasil Jobtailor Tempo integral
  • Serve as the owner of the reliability of Orders, Portability, and Digital Services systems
  • Define and monitor SLOs, SLIs, and Error Budgets
  • Work to meet SLAs and reduce MTTR/MTTD
  • Build proactive automations, runbooks, and AI agents for incident prevention and self-remediation
  • Enhance the metrics, logging, and tracing stack
  • Ensure actionable alerts and system health dashboards
  • Participate in the on-call rotation
  • Lead postmortems with root-cause analysis (RCA) and action plans
  • Support development teams in designing cloud-native applications, disaster recovery (DR), chaos engineering, and load testing
  • Improve deployment pipelines and promote SRE best practices alongside development teams
  • Foster an SRE, blameless, and continuous improvement culture
  • Provide technical leadership to the area’s squads and reduce toil while increasing system resilience and predictability

Requirements

  • 4+ years of experience in SRE, DevOps, or Platform Engineering
  • Experience with cloud environments: AWS, Azure, or GCP
  • Proficiency in Kubernetes, Docker, Terraform, and CI/CD
  • Knowledge of Jenkins, GitLab CI, and Azure DevOps
  • Knowledge of Java 17/21, Spring Boot, REST APIs, and JDBC
  • Knowledge of Oracle, OCI, databases, and application servers
  • Knowledge of Angular
  • Knowledge of microservices, Postman, REST, and caching
  • Knowledge of observability using Prometheus, Grafana, ELK, Datadog, or similar tools
  • Experience with incident management, on-call rotations, and postmortems
  • Knowledge of SLOs, SLIs, and Error Budgets
  • Experience in telecommunications companies is a plus
  • Knowledge of Orders, Portability, and Digital Services systems is a plus
  • Experience with AI agents and proactive automation is a plus
  • Data-driven mindset with a results-oriented approach
  • Clear communication skills for managing crises and engaging stakeholders
  • Proactive approach to identifying bottlenecks before they become incidents
  • Collaborative work with teams in São Paulo and Curitiba

Core Competencies

Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on cloud-native application design, incident management, and proactive automation. Capable of defining and monitoring SLOs, SLIs, and Error Budgets while fostering a culture of continuous improvement and collaboration.

Highest-signal resume keywords

  • Site Reliability Engineering (SRE)
  • Cloud Environments: AWS, Azure, GCP
  • Kubernetes, Docker, Terraform
  • Incident Management and Postmortems
  • Proactive Automation and AI Agents

ATS Optimization Keywords

Hard Skills

  • SLOs, SLIs, and Error Budgets
  • Java 17/21, Spring Boot
  • REST APIs and JDBC
  • Microservices and Caching
  • Observability: Prometheus, Grafana, ELK, Datadog

Soft Skills

  • Clear Communication Skills
  • Proactive Approach
  • Collaborative Work

Industry Keywords

  • Telecommunications
  • Orders, Portability, and Digital Services Systems

Tools & Technologies

  • Jenkins
  • GitLab CI
  • Azure DevOps
  • Postman