Amazon · amazon.jobs · checked today
Sr. Data Engineer, Amazon Digital Advertising
Application deadline: Sep 26, 2026 Are you excited by the idea of building the data foundation that an entire AI-powered product is built on — from the very first pipeline? Do you like the messy, ambiguous problems: wrangling inconsistent data from dozens of outside partners and turning it into something clean, timely, and genuinely trustworthy?
Skills, with evidence
- Data modelling
Experience with data modeling, warehousing and building ETL pipelines
must have - Python
Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
must have - SQL
Experience with SQL
must have - Spark
Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
must have - Warehousing
Experience with data modeling, warehousing and building ETL pipelines
must have - Data quality
You'll figure out how to reliably pull data from noisy, ever-changing external sources, reconcile feeds that never quite agree with each other, and build the quality checks that catch problems before anyone downstream ever sees them.
- Airflow / orchestration
- Partner with modeling and AI engineers to ensure the data foundation is structured for downstream analytics, agent orchestration, and natural-language query.
- Failure handling
- Instrument pipelines for observability, freshness/SLA monitoring, and low-touch operations, so the service requires primarily configuration updates as new sources are added.
- Java
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
not practised here - Scala
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
SQL · Data quality
- Median delivery time per cityIntermediate
- Bucket deliveries into quartilesIntermediate
- Median order value without a median functionIntermediate
- New and repeat orders by monthIntermediate
- Every order against its customer's averageIntermediate
- Python: the data-wrangling round≈ 3 h
Python
- Diff two snapshots of a tableIntermediate
- Explode an array column into rowsIntermediate
- Flatten nested event payloadsIntermediate
- Pivot a long metrics table to wideIntermediate
- Choose what an incremental run should readIntermediate
- Data modelling: the round most people fail≈ 3 h
Data modelling · Warehousing
- Addresses that stay true to the pastIntermediate
- Seat holds and the release-night raceIntermediate
- Subscription warehouse grainIntermediate
- Campaign efficiencyIntermediate
- Catalogue: products, variants and sellersIntermediate
- Pipeline design: safe to run twice≈ 4 h
Airflow / orchestration · Failure handling
- Is this change safe?Intermediate
- Marketplace transactions at scaleIntermediate
- The source will not let youIntermediate
- Changing a pipeline that’s already runningIntermediate
- SLA-aware alerting flowIntermediate
- Spark: read the plan Spark actually ran≈ 2 h
Spark
- broadcast() with auto-broadcast off, and the case where Spark ignores itIntermediate
- autoBroadcastJoinThreshold compares an estimate: flip a join with select() and one settingIntermediate
- Does Spark really run your EXISTS subquery once per row?Intermediate
- left_semi and left_anti: "customers who did / never did" without a full joinIntermediate
- A self-join on a real key that still multiplies rowsIntermediate
- Say it out loud≈ 1 h
25 drills · Intermediate + Advanced≈ 14 hours
Not covered by the plan: Java, Scala.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 6 dremoved the moment Amazon closes it
