Expert DevOps Engineer
Há 2 semanas
Brasil
Ciklum
Remoto
Tempo integral
Grátis com e-mail ou Google
Salve esta vaga e mantenha sua pesquisa organizada
Crie uma conta gratuita para salvar vagas, criar alertas e retornar a esta listagem a partir do seu painel.
Grátis com e-mail ou Google
Ao continuar, você concorda com nossos Termos & Política de Privacidade.
Ciklum is looking for an Expert DevOps Engineer to join our team in Brazil.
We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.
About the role
As an Expert DevOps Engineer, you'll become a part of a cross-functional development team engineering experience of tomorrow.
Responsibilities:
• Platform Strategy & Architecture
• Own and execute the Platform roadmap: compute, networking, identity, observability, shared services, and AI/ML tooling across AWS and Azure
• Lead cloud modernization against the AWS and Azure Well-Architected Frameworks across all five pillars: operational excellence, security, reliability, performance efficiency, and cost optimization
• Define golden paths
- standardized self-service workflows for service scaffolding, DB provisioning, environment spin-up, and AI workload deployment
- with escape hatches for edge cases
• Own multi-cloud strategy; ensure consistent IAM, networking, and FinOps governance across providers
• IaC & CI/CD Automation
• Drive OpenTofu/Ansible as source of truth for all infrastructure; enforce GitOps and policy-as-code for governance, auditability, and security
• Build and mature CI/CD pipelines (GitHub Actions, ArgoCD) to maximize deployment frequency, reduce lead time, and enable zero-ticket self-service provisioning
• Observability
• Own org-wide observability: metrics, logs, traces, and alerting – extended to AI/LLM signals (token usage, model latency, inference cost, agent task completion rates)
• Operate a centralized observability platform (Datadog/Signoz, OpenTelemetry, Grafana/Prometheus/Loki, or equivalent) delivered via golden paths; define SLIs/SLOs as onboarding defaults for all services
• Ensure full-stack coverage across infrastructure, Kubernetes, APM, distributed tracing, AI pipelines, and cost anomaly detection
• Shared Services
• Build and operate a self-service shared services catalog: secrets management, API gateways, model registries, and LLM gateways
• Rationalize duplicative per-team infrastructure; maintain shared services to production SLA standards with clear ownership and consistent security controls
• AI Platform & Agentic Infrastructure
• Own GPU/accelerated compute, model serving, vector databases, RAG pipelines, and LLM API gateway management (AWS Bedrock, Azure OpenAI, Anthropic)
• Build AI golden paths for self-service model deployment and LLM integration; design agentic infrastructure including orchestration runtimes, tool registries, memory/state services, and human-in-the-loop workflows
• Establish governance, cost controls, prompt injection guardrails, and model access policies for AI API usage and inference spend
• Partner with data science and ML engineering to translate agentic workflow requirements into reusable platform primitives
• Platform Adoption & Team Migration
• Collaborate on migration program: partner with peer managers to plan and execute structured workload migrations onto the platform with hands-on support
- not just documentation
• Define onboarding playbooks covering golden paths, shared services, observability setup, CI/CD cutover, and AI capability onboarding; track and report adoption metrics to leadership
• Identify and remove migration blockers
- technical gaps, missing services, or organizational friction — and feed them into the platform roadmap
• Developer Experience, Leadership & Culture
• Build a self-service developer portal (Backstage, GitHub or equivalent) with service catalogs, golden paths, and AI/agentic workflow templates; track DORA metrics and developer experience KPIs
• Hire, develop, and retain high-performing platform engineers; build AI fluency across the team and foster a platform
- as-a-product culture with feedback loops, OKRs, and iterative roadmapping
• Lead architecture reviews; make pragmatic build-vs-buy decisions; partner with security and compliance on governance priorities
• Security, Compliance & FinOps
• Embed secure-by-default guardrails: IaC scanning, RBAC, secrets management, container hardening, and AI-specific controls (prompt injection defense, model access governance, data residency)
• Own cloud cost optimization across AWS and Azure including AI inference spend; maintain SOC 2/ISO 27001 compliance posture
Requirements:
• 8+ years in infrastructure, DevOps, or platform engineering; 2+ years in engineering management
• Cloud: Deep hands-on AWS and Azure expertise: multi-cloud architecture, IAM, networking, compute, and AI/ML services (SageMaker, Bedrock, Azure OpenAI, Azure ML)
• IaC & CI/CD: Terraform required; GitOps, policy-as-code; GitHub Actions / ArgoCD at scale
• DP: Proven track record building an IDP with self-service workflows, golden paths, and developer portal (Backstage, GitHub, or equivalent)
• Observability: OpenTelemetry, Datadog, Signoz, or Prometheus/Grafana at scale; SLI/SLO definition and enforcement
• Shared Services: Built and operated multi-team shared service catalogs with production-grade SLAs
• Adoption: Led structured platform migration and adoption programs in partnership with peer engineering leaders
• Kubernetes & WAF: Kubernetes cluster management, Helm, RBAC, service mesh; AWS and Azure Well-Architected Framework reviews
• Strong cross-functional influencing skills; comfortable as a peer to engineering managers and product leaders Desirable:
• AWS SA Pro / Azure Expert / CKA/CKAD | Python, Go, or Bash What’s in it for you?
• Care: your mental and physical health is our priority. We ensure comprehensive company-paid medical insurance and mental health programs, 5 undocumented sick-leave days per year
• Tailored education path: boost your skills and knowledge with our regular internal events (meetups, conferences, workshops), Udemy license, language courses and company-paid certifications
• Growth environment: share your experience and level up your expertise with a community of skilled professionals, locally and globally
• Long-term employment with 20 working-days paid vacation and local bank holidays
• Flexibility: 100% remote work mode
• Opportunities: we value our specialists and always find the best options for them. Our Internal Mobility Program helps change a project if needed to help you grow, excel professionally and fulfill your potential
• Global impact: work on large-scale projects that redefine industries with international and fast-growing clients
• Welcoming environment: feel empowered with a friendly team, open-door policy, informal atmosphere within the company and regular team-building events
About us
At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress. As we expand into Latin America, every Ciklumer is helping to shape our story. Collaborate with seasoned experts and make a global impact backed by two decades of industry leadership. Explore, empower, engineer with Ciklum Interested already? We would love to get to know you Submit your application. We can’t wait to see you at Ciklum. #LI-IK1
About the role
As an Expert DevOps Engineer, you'll become a part of a cross-functional development team engineering experience of tomorrow.
Responsibilities:
• Platform Strategy & Architecture
• Own and execute the Platform roadmap: compute, networking, identity, observability, shared services, and AI/ML tooling across AWS and Azure
• Lead cloud modernization against the AWS and Azure Well-Architected Frameworks across all five pillars: operational excellence, security, reliability, performance efficiency, and cost optimization
• Define golden paths
- standardized self-service workflows for service scaffolding, DB provisioning, environment spin-up, and AI workload deployment
- with escape hatches for edge cases
• Own multi-cloud strategy; ensure consistent IAM, networking, and FinOps governance across providers
• IaC & CI/CD Automation
• Drive OpenTofu/Ansible as source of truth for all infrastructure; enforce GitOps and policy-as-code for governance, auditability, and security
• Build and mature CI/CD pipelines (GitHub Actions, ArgoCD) to maximize deployment frequency, reduce lead time, and enable zero-ticket self-service provisioning
• Observability
• Own org-wide observability: metrics, logs, traces, and alerting – extended to AI/LLM signals (token usage, model latency, inference cost, agent task completion rates)
• Operate a centralized observability platform (Datadog/Signoz, OpenTelemetry, Grafana/Prometheus/Loki, or equivalent) delivered via golden paths; define SLIs/SLOs as onboarding defaults for all services
• Ensure full-stack coverage across infrastructure, Kubernetes, APM, distributed tracing, AI pipelines, and cost anomaly detection
• Shared Services
• Build and operate a self-service shared services catalog: secrets management, API gateways, model registries, and LLM gateways
• Rationalize duplicative per-team infrastructure; maintain shared services to production SLA standards with clear ownership and consistent security controls
• AI Platform & Agentic Infrastructure
• Own GPU/accelerated compute, model serving, vector databases, RAG pipelines, and LLM API gateway management (AWS Bedrock, Azure OpenAI, Anthropic)
• Build AI golden paths for self-service model deployment and LLM integration; design agentic infrastructure including orchestration runtimes, tool registries, memory/state services, and human-in-the-loop workflows
• Establish governance, cost controls, prompt injection guardrails, and model access policies for AI API usage and inference spend
• Partner with data science and ML engineering to translate agentic workflow requirements into reusable platform primitives
• Platform Adoption & Team Migration
• Collaborate on migration program: partner with peer managers to plan and execute structured workload migrations onto the platform with hands-on support
- not just documentation
• Define onboarding playbooks covering golden paths, shared services, observability setup, CI/CD cutover, and AI capability onboarding; track and report adoption metrics to leadership
• Identify and remove migration blockers
- technical gaps, missing services, or organizational friction — and feed them into the platform roadmap
• Developer Experience, Leadership & Culture
• Build a self-service developer portal (Backstage, GitHub or equivalent) with service catalogs, golden paths, and AI/agentic workflow templates; track DORA metrics and developer experience KPIs
• Hire, develop, and retain high-performing platform engineers; build AI fluency across the team and foster a platform
- as-a-product culture with feedback loops, OKRs, and iterative roadmapping
• Lead architecture reviews; make pragmatic build-vs-buy decisions; partner with security and compliance on governance priorities
• Security, Compliance & FinOps
• Embed secure-by-default guardrails: IaC scanning, RBAC, secrets management, container hardening, and AI-specific controls (prompt injection defense, model access governance, data residency)
• Own cloud cost optimization across AWS and Azure including AI inference spend; maintain SOC 2/ISO 27001 compliance posture
Requirements:
• 8+ years in infrastructure, DevOps, or platform engineering; 2+ years in engineering management
• Cloud: Deep hands-on AWS and Azure expertise: multi-cloud architecture, IAM, networking, compute, and AI/ML services (SageMaker, Bedrock, Azure OpenAI, Azure ML)
• IaC & CI/CD: Terraform required; GitOps, policy-as-code; GitHub Actions / ArgoCD at scale
• DP: Proven track record building an IDP with self-service workflows, golden paths, and developer portal (Backstage, GitHub, or equivalent)
• Observability: OpenTelemetry, Datadog, Signoz, or Prometheus/Grafana at scale; SLI/SLO definition and enforcement
• Shared Services: Built and operated multi-team shared service catalogs with production-grade SLAs
• Adoption: Led structured platform migration and adoption programs in partnership with peer engineering leaders
• Kubernetes & WAF: Kubernetes cluster management, Helm, RBAC, service mesh; AWS and Azure Well-Architected Framework reviews
• Strong cross-functional influencing skills; comfortable as a peer to engineering managers and product leaders Desirable:
• AWS SA Pro / Azure Expert / CKA/CKAD | Python, Go, or Bash What’s in it for you?
• Care: your mental and physical health is our priority. We ensure comprehensive company-paid medical insurance and mental health programs, 5 undocumented sick-leave days per year
• Tailored education path: boost your skills and knowledge with our regular internal events (meetups, conferences, workshops), Udemy license, language courses and company-paid certifications
• Growth environment: share your experience and level up your expertise with a community of skilled professionals, locally and globally
• Long-term employment with 20 working-days paid vacation and local bank holidays
• Flexibility: 100% remote work mode
• Opportunities: we value our specialists and always find the best options for them. Our Internal Mobility Program helps change a project if needed to help you grow, excel professionally and fulfill your potential
• Global impact: work on large-scale projects that redefine industries with international and fast-growing clients
• Welcoming environment: feel empowered with a friendly team, open-door policy, informal atmosphere within the company and regular team-building events
About us
At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress. As we expand into Latin America, every Ciklumer is helping to shape our story. Collaborate with seasoned experts and make a global impact backed by two decades of industry leadership. Explore, empower, engineer with Ciklum Interested already? We would love to get to know you Submit your application. We can’t wait to see you at Ciklum. #LI-IK1