Amazon · amazon.jobs · checked today
Data Engineer I, DSP Analytics
Do you enjoy diving deep into data, developing real-time and batch pipelines that generate actionable insights? The DSP Analytics team has an exciting opportunity for a Data Engineer to make impactful contributions to Amazon's Delivery Service Partner (DSP) program tackling modern data challenges by combining traditional engineering practices with transformative analytics and Generative AI. You'll help design and build the next generation of data infrastructure powering GenAI applications, and develop agents that automate the end-to-end data operations lifecycle.
Skills, with evidence
- Data modelling
Experience with data modeling, warehousing and building ETL pipelines
must have - Python
Experience with one or more scripting language (e.g., Python, KornShell)
must have - SQL
Experience with SQL
must have - Spark
Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
must have - Warehousing
Experience with data modeling, warehousing and building ETL pipelines
must have - AWS
Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
must have · not practised here - Cost & performance
Develop and optimize ETL (Extract, Transform, Load) processes to clean, enrich, and structure raw data into a usable format for analysis and reporting.
- Streaming
Do you enjoy diving deep into data, developing real-time and batch pipelines that generate actionable insights?
- Data quality
Establish data quality standards, perform data validation, and proactively identify and address data quality issues.
- Failure handling
Implement monitoring solutions to proactively detect and address data pipeline failures or performance bottlenecks.
- Governance & security
Ensure data privacy and security by implementing access controls, encryption, and compliance with data protection regulations.
not practised here - Machine learning
We are a team of Data Engineers working closely with Data Scientists, Economists, and Analysts turning machine learning and AI research into scalable products that delight customers worldwide.
not practised here - Scala
- Experience with one or more query language (e.g., SQL, PL/SQL, DDL, MDX, HiveQL, SparkSQL, Scala)
not practised here
Your plan
- The SQL screen: correct, then fast≈ 1 h
SQL · Data quality
- Average order value by countryFoundations
- Line items per orderFoundations
- Revenue by buyer countryFoundations
- Customers who never orderedFoundations
- Hiring funnel by roleFoundations
- Python: the data-wrangling round≈ 2 h
Python
- Clean spreadsheet headers into column namesFoundations
- Events normalization jobFoundations
- List the distinct composite keysFoundations
- Deduplicate rows, first one winsFoundations
- Render a byte count for humansFoundations
- Data modelling: the round most people fail≈ 2 h
Data modelling · Warehousing
- Marketplace core entitiesFoundations
- Cinema seat bookingFoundations
- City parking baysFoundations
- Dating app matchesFoundations
- Food delivery ordersFoundations
- Pipeline design: safe to run twice≈ 3 h
Streaming · Failure handling
- After the first one finishesFoundations
- The history nobody keptFoundations
- The spreadsheet is a dependencyFoundations
- Waiting for the fileFoundations
- Where did the rows go?Foundations
- Spark: read the plan Spark actually ran≈ 1 h
Spark · Cost & performance
- HAVING vs WHERE: where does a filter after groupBy actually run?Foundations
- countDistinct vs approx_count_distinct: what the extra shuffle buysFoundations
- COUNT(*) vs COUNT(column): the null trapFoundations
- Grouping by two columns: what changes in the shuffle?Foundations
- Does the join type change the join strategy?Foundations
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Governance & security, Machine learning, Scala.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 7 dremoved the moment Amazon closes it
