Site Reliability Engineer

Há 3 dias


São Bernardo do Campo, Brasil Agileengine Tempo inteiro

Site Reliability Engineer (Middle/Senior) ID*****4 weeks ago Be among the first 25 applicantsAgileEngine is an Inc. **** company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries.We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet youWHAT YOU WILL DOShift: Monday – Thursday 8AM – 7PM PST (11AM – 10PM EST) with rotating on-call;On call shifts: every 6 weeks, for one week as primary responder and next week as secondary;Manage alerts daily, check systems, and elevate issues as needed;Be part of a team that provides 24×7 on-call support for critical SaaS events;Be available in case of emergencies when team members are not available or need help;Document issues and remediation steps;Proactively create appropriate monitors in the EKS/K8S ecosystem;Deploy to EKS/K8s cluster using Terraform and Helm;Learn and maintain existing infrastructure running under Docker Swarm;Improve existing infrastructure health by implementing checks and scripts to correct known issues;Maintain and develop deployment code;Automate manual tasks;Implement/integrate new technologies in our Cloud Infrastructure;Collaborate with other teams and departments to provide the highest level of support and assistance;Apply a real customer focus when planning deployments/updates, having the customer in the forefront of the mind, and considering the impact on them before making changes;Work closely on solutions with Support, Customer Success, Migration, and Professional Services teams to provide the best in class SaaS service to our customers;Perform RCA and take necessary corrective actions to prevent the recurrence of issues;Create and assign alert-related actions to the appropriate team after the investigation;Handle support requests for environment-specific actions;Identify and provide automation requirements to improve RCA.MUST HAVES2+ years of professional experience;Experience working with Datadog;Hands-on experience as an AWS Cloud Engineer;Working knowledge of EKS/Terraform/Helm;Working Experience with Docker and Docker Swarm;Good understanding of AWS IAM roles and policies;Experience logging and monitoring AWS resources using CloudWatch logs;Experience working in a Linux environment;Proficient in Bash and/or Python scripting;A strong understanding of web technologies such as REST APIs;Working Experience with monitoring solutions, such as Grafana and Prometheus;Excellent oral and written communication skills;Customer-facing communication skills to effectively explain issues and RCAs to them;Experience in Product/Application Support for SaaS-based products;Understanding of APIs, Databases, Systems Architecture, and Design;Designing, implementing, and operating in a DevSecOps;Excellent communication skills, both written and verbal;Ability to work independently as well as within a collaborative environment;A technical aptitude with the desire to learn new and evolving technologies;Upper-Intermediate English level.NICE TO HAVESExperience with GCP or Azure;Certifications: AWS Certified DevOps Engineer – Professional or AWS Certified Advanced Networking Specialty.PERKS AND BENEFITSProfessional growth: Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps.Competitive compensation: We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities.A selection of exciting projects: Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands.Flextime: Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.Seniority levelMid-Senior levelEmployment typeFull-timeJob functionIT Services and IT ConsultingReferrals increase your chances of interviewing at AgileEngine by 2xWe're unlocking community knowledge in a new way.Experts add insights directly into each article, started with the help of AI.#J-*****-Ljbffr


  • Site Reliability Engineer

    2 semanas atrás


    São Bernardo do Campo, Brasil BairesDev Tempo inteiro

    OverviewSite Reliability Engineer - Remote Work | REF# Join to apply for the Site Reliability Engineer - Remote Work role at BairesDev. We are looking for a Site Reliability Engineer to administrate and provide support for the whole project infrastructure hosted in the cloud while implementing CI/CD pipelines for the automation of the deployments....

  • Senior Site Reliability

    3 semanas atrás


    Campo Grande, Brasil Canonical Tempo inteiro

    Senior Site Reliability / Gitops Engineer Join to apply for the Senior Site Reliability / Gitops Engineer role at Canonical Senior Site Reliability / Gitops Engineer 1 day ago Be among the first 25 applicants Join to apply for the Senior Site Reliability / Gitops Engineer role at Canonical Get AI-powered advice on this job and more exclusive features....


  • São Paulo, Brasil K2 Solutions Tempo inteiro

    Trabalho híbrido na região de Pinheiros/ SP - 3x por semana no escritórioEstamos selecionando um Senior Site Reliability Engineer - SRE para se juntar ao nosso time e desempenhar um papel essencial na manutenção, automação e melhoria da confiabilidade dos sistemas que impulsionam a rede logística da empresa em múltiplas regiões. Essa pessoa...

  • Site Reliability Engineer

    4 semanas atrás


    Campo Grande, Brasil INDI Staffing Services Tempo inteiro

    Overview We are looking for a Site Reliability Engineer to build and maintain highly reliable, scalable, and secure OpenShift/Kubernetes clusters. We will need you to approach the problem of building and maintaining production systems from a software engineering perspective with a focus on automation, and reliability. Responsibilities Build, automate, and...


  • São Paulo, Brasil INDI Staffing Services Tempo inteiro

    OverviewWe are looking for a Site Reliability Engineer to build and maintain highly reliable, scalable, and secure OpenShift/Kubernetes clusters. Approach the problem of building and maintaining production systems from a software engineering perspective with a focus on automation and reliability. ResponsibilitiesBuild, automate, and maintain...

  • Site Reliability Engineer

    4 semanas atrás


    São Paulo, Brasil INDI Staffing Services Tempo inteiro

    At INDI, we're passionate about empowering individuals and businesses worldwide. Our cutting-edge recruiters connect leading companies with top talent, fostering a dynamic environment where innovation thrives. Join us in shaping the future of work.Overview of the role:We are looking for a Site Reliability Engineer to build and maintain highly reliable,...

  • Site Reliability Engineer

    3 semanas atrás


    São Paulo, Brasil INDI Staffing Services Tempo inteiro

    At INDI, we're passionate about empowering individuals and businesses worldwide. Our cutting-edge recruiters connect leading companies with top talent, fostering a dynamic environment where innovation thrives. Join us in shaping the future of work. Overview of the role: We are looking for a Site Reliability Engineer to build and maintain highly reliable,...


  • Jaraguá do Sul, Brasil INDI Staffing Services Tempo inteiro

    At INDI, we're passionate about empowering individuals and businesses worldwide. Our cutting-edge recruiters connect leading companies with top talent, fostering a dynamic environment where innovation thrives. Join us in shaping the future of work. Overview of the role: We are looking for a Site Reliability Engineer to build and maintain highly reliable,...


  • São Paulo, Brasil Chainlink Labs Tempo inteiro

    Join to apply for the Senior Site Reliability Engineer role at Chainlink Labs 2 weeks ago Be among the first 25 applicants Join to apply for the Senior Site Reliability Engineer role at Chainlink Labs Get AI-powered advice on this job and more exclusive features. About UsChainlink Labs is the primary contributing developer of Chainlink, the decentralized...


  • São Paulo, Brasil INDI Staffing Services Tempo inteiro

    Overview of the role We are looking for a Site Reliability Engineer to build and maintain highly reliable, scalable, and secure OpenShift/Kubernetes clusters. We will need you to approach the problem of building and maintaining production systems from a software engineering perspective with a focus on automation, and reliability. Key responsibilities Build...