Data Engineer

Raising the Village
Raising the Village

Software Engineering, Data Science

Mbarara, Uganda

Posted on Aug 25, 2026

Job Title: Data Engineer

Department/Group: Venn

Reporting To: Senior Data Scientist

Years of Experience: 1-3 years

Location: Mbarara

Travel Required: 20% Job Description

About Raising The Village

We are Raising The Village (RTV) – an international development organization and a registered charity – on a mission to end ultra-poverty in sub-Saharan Africa. Raising The Village is a fast-growing organization on an accelerated growth path. We have 350+ national staff in the Sub-Saharan Africa (SSA) region and a team of 15+ people in North America working together to lift communities out of ultra-poverty in last-mile villages. We operate at the intersection of direct implementation and advanced data analytics to inform progress, decision-making, and impact.

To date, we have supported more than 1,000,000 people in SSA through our innovative holistic approach and are on track to expand our reach and impact year over year.

We have achieved this tremendous growth with the support of our incredible partners from all around the globe who believe in our model and impact. Find out more about our programs and impact at: www.raisingthevillage.org.

The VENN department is the data and technology backbone of our organization, connecting advanced analytics and custom software tools with field implementation to ensure data-informed decision-making at every level.

Role Description

The Data Engineer is a core builder on RTV's expanding data infrastructure team, reporting directly to the Senior Data Scientist within the VENN department, and responsible for designing, developing, and maintaining the pipelines, warehouse layers, and data quality systems that power RTV's programmatic analytics, machine learning platforms, and field evaluation tools. Operating at the intersection of data platform engineering, ML infrastructure support, and field data integration, this role works across a fast-moving roadmap that spans batch and streaming ingestion, ELT pipeline development, Delta Lake architecture, observability frameworks, and the integration of structured field data with AI model outputs.

Key Responsibilities

Pipeline Development & Delivery

●Design and deliver batch and streaming data pipelines that ingest data from field collection platforms (SurveyCTO, ArcGIS, custom mobile apps) into RTV's Databricks warehouse, working at pace against an accelerated infrastructure roadmap.

●Implement ELT and ETL workflows in PySpark and Databricks SQL, applying Medallion Architecture (Bronze, Silver, Gold) principles to produce clean, versioned, and consumption-ready data layers.

●Contribute actively to sprint-based delivery cycles, taking ownership of pipeline workstreams end-to-end from design through deployment and monitoring.

Delta Lake & Warehouse Architecture

●Build and maintain Delta Lake table structures with strong schema enforcement, ACID-compliant write patterns, and time travel capabilities to support auditability across program datasets.

●Contribute to data modeling decisions including; star schema design, SCD patterns, and denormalization trade-offs.

●Support the evolution of Unity Catalog governance structures including lineage tracking, access controls, and dataset documentation as the warehouse scales across new program domains and geographies.

Data Observability & Quality

●Implement and maintain data observability frameworks across all pipeline stages, covering data freshness, volume anomalies, schema drift detection, and SLA monitoring.

●Build validation and quality checks at ingestion and transformation layers using tools such as Great Expectations, dbt tests, or Databricks-native monitoring capabilities.

●Establish structured logging, alerting, and incident response practices that give the team fast, reliable visibility into pipeline health across all environments.

ML & AI Pipeline Support

●Integrate structured household data, image classification outputs, and ML model predictions into unified warehouse layers for consumption by Data Scientists and the WorkMate AI platform.

●Collaborate with ML Engineers and Data Scientists to build and maintain feature engineering pipelines and training data preparation workflows that support RTV's computer vision and adoption scoring systems.

Collaboration & Documentation

●Work closely with the Senior Data Engineer on roadmap prioritization, architectural decisions, and engineering standards as the team scales.

●Partner with Software Engineers, Data Scientists, field evaluation teams, and program staff to understand data requirements and translate them into reliable, well-documented pipeline solutions.

●Maintain thorough documentation of pipeline architectures, transformation logic, data dictionaries, and runbooks to support team growth and organizational knowledge continuity.

Technical Requirements

Education & Experience

•A Bachelor's degree in Computer Science, Software Engineering, Data Engineering, Information Systems, Statistics, or a related quantitative field is preferred.

•Equivalent practical experience through demonstrable project work, open-source contributions, or bootcamp training is equally welcome.

•Clear evidence of building and shipping production-grade pipelines, including demonstrable ownership from design through deployment with specific examples of integrating, moving, and transforming data at a meaningful scale.

Technical Skills

•Candidates must demonstrate proficiency in Python (including Pandas), PySpark, and advanced SQL for building and transforming data at scale within a Databricks-first environment spanning Delta Lake, Unity Catalog, and Medallion Architecture.

•Core engineering competencies include ELT/ETL design, batch and Spark Structured Streaming pipelines, and data observability practices covering validation, monitoring, and alerting frameworks. Supporting skills include Git-based collaborative workflows, working knowledge of AWS core services, and the ability to produce clean, visualization-ready datasets for consumption in tools such as Power BI, Tableau, or Python visualization libraries.