SRE Consultant – Site Reliability Engineering
Há 7 dias
São Paulo, State of São Paulo, Brasil
Jobtailor
Tempo integral
Grátis com e-mail ou Google
Salve esta vaga e mantenha sua pesquisa organizada
Crie uma conta gratuita para salvar vagas, criar alertas e retornar a esta listagem a partir do seu painel.
Grátis com e-mail ou Google
Ao continuar, você concorda com nossos Termos & Política de Privacidade.
- Serve as the owner of the reliability of Orders, Portability, and Digital Services systems
- Define and monitor SLOs, SLIs, and Error Budgets
- Work to meet SLAs and reduce MTTR/MTTD
- Build proactive automations, runbooks, and AI agents for incident prevention and self-remediation
- Enhance the metrics, logging, and tracing stack
- Ensure actionable alerts and system health dashboards
- Participate in the on-call rotation
- Lead postmortems with root-cause analysis (RCA) and action plans
- Support development teams in designing cloud-native applications, disaster recovery (DR), chaos engineering, and load testing
- Improve deployment pipelines and promote SRE best practices alongside development teams
- Foster an SRE, blameless, and continuous improvement culture
- Provide technical leadership to the area’s squads and reduce toil while increasing system resilience and predictability
Requirements
- 4+ years of experience in SRE, DevOps, or Platform Engineering
- Experience with cloud environments: AWS, Azure, or GCP
- Proficiency in Kubernetes, Docker, Terraform, and CI/CD
- Knowledge of Jenkins, GitLab CI, and Azure DevOps
- Knowledge of Java 17/21, Spring Boot, REST APIs, and JDBC
- Knowledge of Oracle, OCI, databases, and application servers
- Knowledge of Angular
- Knowledge of microservices, Postman, REST, and caching
- Knowledge of observability using Prometheus, Grafana, ELK, Datadog, or similar tools
- Experience with incident management, on-call rotations, and postmortems
- Knowledge of SLOs, SLIs, and Error Budgets
- Experience in telecommunications companies is a plus
- Knowledge of Orders, Portability, and Digital Services systems is a plus
- Experience with AI agents and proactive automation is a plus
- Data-driven mindset with a results-oriented approach
- Clear communication skills for managing crises and engaging stakeholders
- Proactive approach to identifying bottlenecks before they become incidents
- Collaborative work with teams in São Paulo and Curitiba
Core Competencies
Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on cloud-native application design, incident management, and proactive automation. Capable of defining and monitoring SLOs, SLIs, and Error Budgets while fostering a culture of continuous improvement and collaboration.
Highest-signal resume keywords
- Site Reliability Engineering (SRE)
- Cloud Environments: AWS, Azure, GCP
- Kubernetes, Docker, Terraform
- Incident Management and Postmortems
- Proactive Automation and AI Agents
ATS Optimization Keywords
Hard Skills
- SLOs, SLIs, and Error Budgets
- Java 17/21, Spring Boot
- REST APIs and JDBC
- Microservices and Caching
- Observability: Prometheus, Grafana, ELK, Datadog
Soft Skills
- Clear Communication Skills
- Proactive Approach
- Collaborative Work
Industry Keywords
- Telecommunications
- Orders, Portability, and Digital Services Systems
Tools & Technologies
- Jenkins
- GitLab CI
- Azure DevOps
- Postman