Role Overview
We are seeking a Machine Learning Data Engineer to build and operate large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities.
What You Will Do
The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and a clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Why It Might Be a Fit
The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and a clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field
- Six or more years of data engineering experience, with significant work supporting ML or AI workloads
- Strong proficiency in Python and at least one JVM or systems language
- Deep experience with modern data processing frameworks such as Spark, Ray, or Beam
- Hands-on experience operating petabyte-scale storage and pipeline systems
- Strong understanding of distributed systems, data modeling, and storage formats
- Experience with dataset versioning, lineage, and reproducibility for ML workflows
- Familiarity with high-throughput data loading for accelerator-based training
- Strong software engineering practices including testing, CI/CD, and code review
- Excellent communication and cross-functional collaboration skills
Benefits
- Salary Range: $100,000–$150,000 Annually
To apply for this job please visit brightvisiontechnologies.applytojob.com.

Follow us on social media