Amazon · amazon.jobs · checked today
Data Engineer II, DSP Analytics
Do you enjoy diving deep into data, developing real-time and batch pipelines that generate actionable insights? The DSP Analytics team has an exciting opportunity for a Data Engineer to make impactful contributions to Amazon's Delivery Service Partner (DSP) program tackling modern data challenges by combining traditional engineering practices with transformative analytics and Generative AI. You'll help design and build the next generation of data infrastructure powering GenAI applications, and develop agents that automate the end-to-end data operations lifecycle.
Skills, with evidence
- Data modelling
Experience with data modeling, warehousing and building ETL pipelines
must have - Python
Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
must have - Warehousing
Experience with data modeling, warehousing and building ETL pipelines
must have - AWS
Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
must have · not practised here - Governance & security
Establish data engineering best practices, including governance, security standards, and operational excellence
must have · not practised here - SQL
- 2+ years of analyzing and interpreting data with Redshift, Oracle, NoSQL etc.
- Streaming
Do you enjoy diving deep into data, developing real-time and batch pipelines that generate actionable insights?
- Airflow / orchestration
- Familiarity with agentic AI patterns including tool use, function calling, and multi-agent orchestration
- Spark
- Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
- Go
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
not practised here - Java
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
not practised here - Machine learning
We are a team of Data Engineers working closely with Data Scientists, Economists, and Analysts turning machine learning and AI research into scalable products that delight customers worldwide.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 1 h
SQL
- Line items per orderFoundations
- Average order value by countryFoundations
- Count at the right grainFoundations
- Revenue by buyer countryFoundations
- Join three tables into an order detailFoundations
- Python: the data-wrangling round≈ 2 h
Python
- Clean spreadsheet headers into column namesFoundations
- Events normalization jobFoundations
- List the distinct composite keysFoundations
- Deduplicate rows, first one winsFoundations
- Render a byte count for humansFoundations
- Data modelling: the round most people fail≈ 2 h
Data modelling · Warehousing
- Marketplace core entitiesFoundations
- Cinema seat bookingFoundations
- City parking baysFoundations
- Dating app matchesFoundations
- Food delivery ordersFoundations
- Pipeline design: safe to run twice≈ 3 h
Streaming · Airflow / orchestration
- After the first one finishesFoundations
- Run it for last TuesdayFoundations
- The history nobody keptFoundations
- The spreadsheet is a dependencyFoundations
- Too slow by morningFoundations
- Spark: read the plan Spark actually ran≈ 1 h
Spark
- HAVING vs WHERE: where does a filter after groupBy actually run?Foundations
- countDistinct vs approx_count_distinct: what the extra shuffle buysFoundations
- COUNT(*) vs COUNT(column): the null trapFoundations
- Grouping by two columns: what changes in the shuffle?Foundations
- Does the join type change the join strategy?Foundations
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Governance & security, Go, Java, Machine learning.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 7 dremoved the moment Amazon closes it
