Role Overview
Join a diverse, global community working across cultures, disciplines, and borders to address the world’s most pressing development challenges. The role is anchored in data engineering with some elements of product management.
What You Will Do
Design, build, and maintain scalable, secure, and reliable batch and streaming data pipelines using Python, SQL, PySpark, and modern data processing frameworks.
Why It Might Be a Fit
The successful candidate will clarify business intent, shape priorities, define acceptance criteria, make informed trade-offs, and remain accountable for the reliability, adoption, cost, and continuous improvement of assigned data products.
Requirements
- 5+ years of experience building agentic workflows
- Extensive experience building agentic workflows is not required, but candidates should demonstrate sound judgment, active learning, and openness to changing how data products are designed, built, tested, released, and operated.
- Hands-on use of AI-enabled engineering practices
- Ability to show how they already use AI assistants or agents in real engineering work, in their current role or in their own projects, and how they check the results.
- Ability to design, build, and maintain scalable, secure, and reliable batch and streaming data pipelines using Python, SQL, PySpark, and modern data processing frameworks.
- Ability to lead ETL and ELT solutions for structured, semi-structured, and unstructured data, including metadata enrichment and AI-ready preparation for analytics and GenAI use cases.
- Ability to develop logical and physical data models, curated datasets, and lakehouse layers optimized for analytics, reporting, operational use, and downstream consumption.
- Ability to establish and evolve data contracts covering schemas, definitions, ownership, service levels, quality expectations, and change-management arrangements.
- Ability to set technical direction for assigned data products and ensure alignment with approved enterprise architecture patterns and standards.
- Ability to analyze schema changes and downstream impacts, align solutions with master and reference data, and guide decisions on versioning and migration.
- Ability to review designs and code, resolve complex engineering issues, and ensure solutions are reusable, maintainable, and fit for production.
- Ability to modernize legacy data workloads onto the Databricks lakehouse, including analysis of existing logic, reconciliation of old and new outputs, and controlled cutover.
- Ability to apply modern engineering practices on Databricks, including Git-based development, Databricks Asset Bundles or equivalent infrastructure-as-code, Lakeflow pipelines and jobs, and automated testing.
- Ability to translate business needs and intended decisions into clear data product outcomes, requirements, and acceptance criteria.
- Ability to define the product’s users, expected use, service levels, quality thresholds, and measurable indicators of adoption and value.
- Ability to own the data product across its lifecycle, from source discovery and design through release, adoption, operation, improvement, and retirement.
- Ability to make informed trade-offs among scope, value, quality, delivery timing, technical debt, cost, and operational risk.
- Ability to use consumer feedback, usage information, production incidents, quality results, and operating costs to prioritize improvements.
- Ability to communicate product decisions, progress, risks, dependencies, and outcomes clearly to technical and non-technical stakeholders.
- Ability to register and maintain owned data products in the enterprise data product inventory (Collibra), and take them through the IFC data product framework’s review stages.
- Ability to work with business analytics teams who own metric meaning (for example COMAR) to encode agreed business definitions and shared metrics consistently in the semantic layer.
- Ability to clarify the decision each data product serves, identify intended consumers, define success measures, and confirm acceptance criteria before build begins.
- Ability to lead source discovery and acquisition by identifying systems of record, profiling data, documenting grain and source behavior, and coordinating access and data-use approvals.
- Ability to use approved AI tools to support modeling, pipeline development, testing, documentation, release preparation, monitoring, incident analysis, and consumer support where appropriate.
- Ability to determine what work should remain human-led, what can be AI-augmented, and what may be automated within approved controls.
- Ability to critically review and validate AI-generated outputs, retaining accountability for technical quality, security, privacy, data use, and production outcomes.
- Ability to capture source behavior, definitions, contracts, design decisions, quality rules, reviewer corrections, incidents, and root causes as reusable organizational context.
- Ability to experiment frequently and responsibly with emerging AI-enabled engineering practices, evaluate results on live work, and share effective approaches and lessons with colleagues.
- Ability to write specifications, acceptance tests, and instructions clear enough for AI agents to execute, and maintain them as reusable team assets (for example prompt libraries, agent instructions, and skills).
- Ability to make data products usable by AI agents and natural-language tools by maintaining business descriptions, metric definitions, and semantic metadata in Unity Catalog, and by preparing and maintaining spaces for tools such as Databricks Genie.
- Ability to design and run evaluations for AI-generated outputs that touch data products, such as natural-language query answers or AI-based extraction, using reference question sets, expected results, and tracked accuracy over time.
- Ability to optimize data processing through effective partitioning, storage, caching, indexing, orchestration, and workload design.
- Ability to monitor and improve reliability, freshness, service levels, performance, and cost for assigned data products.
- Ability to implement data lineage, classification, access control, auditing, and other governance capabilities using platforms such as Unity Catalog and Collibra.
- Ability to ensure adherence to enterprise standards for privacy, security, responsible AI, risk management, and separation of duties.
- Ability to establish automated quality controls, reconciliation, test coverage, deployment gates, rollback plans, and production-readiness evidence.
- Ability to lead incident triage and root-cause analysis for complex issues, convert lessons into preventive controls, and guide decisions on remediation or retirement.
- Ability to use continuous integration and deployment practices to promote changes safely across environments.
- Ability to use data observability tools (for example Monte Carlo) to monitor freshness, volume, schema, and quality, and connect alerts to clear ownership and response.
- Ability to track and attribute compute and storage cost for owned products, and apply cost controls in line with the team’s FinOps practices.
- Ability to partner with business stakeholders, data owners, data scientists, analysts, architects, security specialists, and platform teams to deliver usable and trusted data products.
- Ability to facilitate decisions on definitions, sourcing, design, quality thresholds, release readiness, and residual risk.
- Ability to provide technical leadership and mentorship to engineers, including design guidance, code review, problem solving, and knowledge sharing.
- Ability to contribute reusable patterns, standards, and practices that improve consistency and delivery across teams.
- Ability to present complex technical choices and recommendations clearly to senior technical and business stakeholders.
- Ability to help colleagues adopt AI-enabled engineering practices through pairing, demonstrations, and shared examples from live work.
Benefits
- Dental insurance
- Vision insurance
- Health insurance
- Paid time off
- Retirement plan
- Learning budget
- Parental leave
- Wellness program
- Visa sponsorship
- Relocation assistance
To apply for this job please visit worldbankgroup.csod.com.

Follow us on social media