Mastercard · mastercard.wd1.myworkdayjobs.com · checked today
Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow)
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible.
Skills, with evidence
- Airflow / orchestration
Experience with workflow orchestration platforms such as Apache Airflow, Databricks Workflows, AWS Step Functions, Azure Data Factory, or similar technologies.
must have - Data modelling
Advanced SQL expertise, including data modeling, performance tuning, query optimization, and large-scale analytical processing.
must have - Data quality
Strong understanding of data governance, metadata management, lineage, quality frameworks, privacy controls, and access management.
must have - Python
Expert programming skills in Python, PySpark, and modern software engineering practices.
must have - SQL
Advanced SQL expertise, including data modeling, performance tuning, query optimization, and large-scale analytical processing.
must have - Spark
Strong hands-on experience with distributed data processing technologies such as Apache Spark and modern lakehouse platforms.
must have - Streaming
Experience building scalable batch and real-time data pipelines.
must have - AWS
Experience building and operating cloud-native data platforms in AWS, or Azure, or other enterprise cloud environments.
must have · not practised here - Governance & security
Implement data governance capabilities including quality controls, lineage, metadata management, cataloging, access control, retention, and compliance.
not practised here - Machine learning
We are seeking a Principal Data Engineer to design and build the data foundations that power advanced analytics, machine learning, generative AI, and agentic AI solutions.
not practised here - Cost & performance
Optimize data workloads for performance, scalability, reliability, resiliency, and cost efficiency.
- Azure
Experience building and operating cloud-native data platforms in AWS, or Azure, or other enterprise cloud environments.
not practised here - Warehousing
Develop governed lakehouse and modern data platform architectures using cloud-native technologies and distributed data processing frameworks.
- Terraform
Experience implementing CI/CD, automated testing, version control, Infrastructure as Code, and platform automation.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
Data quality · SQL
- Airbnb: converting on a day with no published rateAdvanced
- Amazon: net revenue across two fan-outsAdvanced
- Build a percentile tableAdvanced
- Duplicate snapshot keysAdvanced
- Flag outlier ordersAdvanced
- Python: the data-wrangling round≈ 4 h
Python
- Data modelling: the round most people fail≈ 4 h
Data modelling · Warehousing
- Pipeline design: safe to run twice≈ 4 h
Airflow / orchestration · Streaming
- Spark: read the plan Spark actually ran≈ 2 h
Spark · Cost & performance
- applyInPandas per country: what Spark ships to Python, and the native rewriteAdvanced
- Filters you wrote in the wrong place: where Catalyst moves themAdvanced
- snappy, gzip or zstd: measure the Parquet codec trade-off yourselfAdvanced
- A correlated COUNT(*) subquery: the join Spark runs, and the count bug it avoidsAdvanced
- Find the four suspects in a slow pipeline (one of them is innocent)Advanced
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Governance & security, Machine learning, Azure, Terraform.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 12 dremoved the moment Mastercard closes it
