Identity Fraud Detection Pipeline (Real-time)
Designed a real-time fraud detection pipeline ingesting identity events via Kafka, executing PySpark operations, and writing outputs to Delta Lake.
Confidence S. is a Data Engineer with several years of experience in designing and implementing data pipelines and warehouses across various cloud platforms. His core technologies include Python, SQL, and Spark, which he uses to build ELT and ETL pipelines, real-time data ingestion systems, and automated workflows. Confidence is proficient with tools such as Apache Kafka, DBT, and Airflow, and is experienced in managing data warehouses on platforms like BigQuery and AWS Redshift. He ensures robust CI/CD practices using GitHub Actions and Docker. Notable projects include developing a real-time identity fraud detection pipeline using Kafka and PySpark, and leading the migration of enterprise data warehouses to Google BigQuery. He also built a scalable data ingestion platform for LLM training datasets, utilizing AWS S3 and BigQuery. Confidence holds certifications in Google Cloud Professional Data Engineering, Microsoft Fabric Data Engineering, and more. He is well-suited for roles that involve building scalable data solutions and enhancing data quality initiatives in cloud environments.
Designed a real-time fraud detection pipeline ingesting identity events via Kafka, executing PySpark operations, and writing outputs to Delta Lake.
Delivered a scalable data ingestion platform capable of processing various file types stored in AWS S3 into BigQuery using Airflow and Docker.
Lead the migration to Google BigQuery, engineered partitioning and clustering of tables, and constructed Airflow DAGs for automated loads.
Other vetted developers with similar skills and experience

Data Engineer
Lagos · 4+ years experience
14 days risk-free trial · No commitments · We handle contracts and payroll