ML Data Pipeline Engineer

3 semanas atrás


Índio do Brasil Prosigliere Tempo inteiro

We're seeking a Data Pipeline Engineer to own and evolve our exercise recognition training data infrastructure. You'll manage the end-to-end pipeline that collects, synchronizes, validates, and prepares IMU sensor and video data for ML model training.

This role combines systems engineering, data quality automation, and hands-on problem-solving in a production environment.

What You'll Do

Pipeline Operations & Improvement

- Maintain and enhance our multi-source data collection system: IMU sensors (via mobile app) and synchronized video streams from gym-based cameras.
- Improve video capture software robustness, particularly handling network interruptions and operational monitoring.
- Deploy and monitor services in remote Linux environments with appropriate DevOps practices.

Data Quality & Validation

- Evolve our Python-based QC engine that validates data pre- and post-annotation
- Implement checks for IMU-video time synchronization, sensor health, and measurement consistency
- Apply digital signal processing techniques to identify sensor failures, connectivity issues, and measurement irregularities.
- Develop validation logic comparing annotations against sensor data to ensure temporal alignment.

Analysis & Troubleshooting

- Perform ad-hoc analysis on ~1,200+ workout tasks to classify failure modes
- Identify whether issues stem from pipeline bugs, sensor problems, or annotation errors
- Prioritize engineering work based on data quality impact and coordinate with annotation team on fixes

Tooling and Visualization

- Maintain and extend our NextJS UI serving annotators, data scientists, and stakeholders
- Create visualizations (Chart.js) for QC metrics and signal analysis
- Integrate with LabelStudio annotation interface

What You Bring

Required

- Strong Python programming skills, particularly for data processing pipelines
- Experience with time-series data and digital signal processing
- Comfortable working in Linux environments and deploying/monitoring remote services
- Ability to debug complex multi-component systems (sensors, video, networks, sync)
- Data quality mindset: designing validation rules, tracking metrics, investigating anomalies
- SQL/database experience for managing pipeline metadata

Highly Valued

- Video processing experience (RTSP streams, encoding, OCR)
- Working with sensor/IoT data and handling connectivity challenges
- NextJS or modern web frameworks for data tooling
- DevOps practices: containerization, monitoring, logging, alerting
- Experience with annotation pipelines and ML training data workflows
- Background in biomechanics, sports science, or wearable sensors

Tech Stack

- Languages: Python (primary), JavaScript/TypeScript (NextJS UI)
- Data: IMU sensor streams, video (RTSP), time-series analysis, DSP
- Tools: LabelStudio, Chart.js, Linux/bash, OCR libraries
- Infrastructure: Remote deployment, monitoring systems

You'll Thrive Here If You

- Enjoy detective work: diagnosing why data doesn't match expectations
- Balance pragmatism with quality: shipping improvements while maintaining reliability
- Communicate well across technical and non-technical stakeholders
- Can work autonomously in a small, mission-driven team


  • ML Data Pipeline Engineer

    3 semanas atrás


    Brasil Prosigliere Tempo inteiro

    We're seeking a Data Pipeline Engineer to own and evolve our exercise recognition training data infrastructure. You'll manage the end-to-end pipeline that collects, synchronizes, validates, and prepares IMU sensor and video data for ML model training. This role combines systems engineering, data quality automation, and hands-on problem-solving in a...


  • Brasil Prosigliere Tempo inteiro

    We're seeking a Data Pipeline Engineer to own and evolve our exercise recognition training data infrastructure. You'll manage the end-to-end pipeline that collects, synchronizes, validates, and prepares IMU sensor and video data for ML model training. This role combines systems engineering, data quality automation, and hands-on problem-solving in a...

  • Ml data pipeline engineer

    3 semanas atrás


    Brasil Prosigliere Tempo inteiro

    We're seeking a Data Pipeline Engineer to own and evolve our exercise recognition training data infrastructure. You'll manage the end-to-end pipeline that collects, synchronizes, validates, and prepares IMU sensor and video data for ML model training. This role combines systems engineering, data quality automation, and hands-on problem-solving in a...

  • Data Engineer

    Há 3 dias


    Índio do Brasil HeartCentrix Solutions Tempo inteiro

    We are seeking a highly skilled Python Data Engineer with an AI/ML focus to join our client's growing data & analytics team in Brazil. This role is ideal for someone who loves building scalable data pipelines, operationalizing machine learning workflows, and partnering closely with data scientists to bring models into production. You will design, develop,...

  • Data Engineer

    Há 3 dias


    Índio do Brasil HeartCentrix Solutions Tempo inteiro

    We are seeking a highly skilled Python Data Engineer with an AI/ML focus to join our client's growing data & analytics team in Brazil. This role is ideal for someone who loves building scalable data pipelines, operationalizing machine learning workflows, and partnering closely with data scientists to bring models into production.You will design, develop, and...

  • Senior Data Engineer

    3 semanas atrás


    Índio do Brasil Pride Global Tempo inteiro

    We're Hiring: Senior Data Engineer | Remote from Brazil | Fluent English required | Location: Remote – Brazil only Contact: Temporary Are you passionate about building scalable data platforms and cutting-edge MLOps solutions? Do you want to work with a top-tier US company revolutionizing e-commerce and circular fashion? We're looking for a Senior Data...


  • Índio do Brasil Pride Global Tempo inteiro

    We're Hiring: Senior Data Engineer | Remote from Brazil | Fluent English required | Location: Remote – Brazil onlyContact: TemporaryAre you passionate about building scalable data platforms and cutting-edge MLOps solutions? Do you want to work with a top-tier US company revolutionizing e-commerce and circular fashion?We're looking for a Senior Data...


  • Índio do Brasil UPBI Data & AI Tempo inteiro

    A UPBI Data & AI, consultoria especializada em soluções digitais, parceira Microsoft e Databricks, está buscando um(a) Data Engineer com sólida experiência em Databricks e Azure para atuação em projeto internacional estratégico. Modelo: Remoto Regime: PJ Responsabilidades: - Desenvolver e otimizar pipelines de dados utilizando Databricks e...

  • Data Engineer

    Há 2 dias


    Brasil HeartCentrix Solutions Tempo inteiro

    We are seeking a highly skilled Python Data Engineer with an AI/ML focus to join our client’s growing data & analytics team in Brazil. This role is ideal for someone who loves building scalable data pipelines, operationalizing machine learning workflows, and partnering closely with data scientists to bring models into production. You will design, develop,...

  • Machine Learning Engineer

    2 semanas atrás


    Índio do Brasil UST España & Latam Tempo inteiro

    We are still looking for talent... and we would love for you to join our team!For over 25 years, UST has worked alongside the world's best companies to make a real impact through business transformation. Driven by technology, inspired by people, and guided by our purpose, UST supports clients from design to implementation. Together, with more than 30,000...