Amazon · amazon.jobs · checked today
Data Engineer II, Business Data Technologies
Amazon’s eCommerce Foundation (eCF) organization is responsible for the core components that drive the Amazon website and customer experience. Serving millions of customer page views and orders per day, eCF builds for scale. As an organization within eCF, the Business Data Technologies (BDT) group is no exception.
Skills, with evidence
- Data modelling
Experience with data modeling, warehousing and building ETL pipelines
must have - Spark
using Amazon Web Service’s (AWS) Redshift, Hive, and Spark
must have - Warehousing
Experience with data modeling, warehousing and building ETL pipelines
must have - AWS
Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
must have · not practised here - SQL
We provide interfaces for our internal customers to access and query the data hundreds of thousands of times per day, using Amazon Web Service’s (AWS) Redshift, Hive, and Spark.
- Streaming
Implement data ingestion routines both real time and batch using best practices in data modeling, ETL/ELT processes leveraging AWS technologies and Big data tools.
- Governance & security
We collect petabytes of data from thousands of data sources inside and outside Amazon including the Amazon catalog system, inventory system, customer order system, page views on the website and Alexa systems.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 1 h
SQL
- Line items per orderFoundations
- Average order value by countryFoundations
- Count at the right grainFoundations
- Revenue by buyer countryFoundations
- Join three tables into an order detailFoundations
- Data modelling: the round most people fail≈ 2 h
Data modelling · Warehousing
- Marketplace core entitiesFoundations
- Cinema seat bookingFoundations
- City parking baysFoundations
- Dating app matchesFoundations
- Food delivery ordersFoundations
- Pipeline design: safe to run twice≈ 4 h
Streaming
- Which day does it belong to?Foundations
- Parcel tracking pipelineIntermediate
- Five minutes behind the sourceIntermediate
- The source will not let youIntermediate
- The fact arrived firstIntermediate
- Spark: read the plan Spark actually ran≈ 1 h
Spark
- HAVING vs WHERE: where does a filter after groupBy actually run?Foundations
- countDistinct vs approx_count_distinct: what the extra shuffle buysFoundations
- COUNT(*) vs COUNT(column): the null trapFoundations
- Grouping by two columns: what changes in the shuffle?Foundations
- Does the join type change the join strategy?Foundations
- Say it out loud≈ 1 h
20 drills · Foundations + Intermediate≈ 9 hours
Not covered by the plan: AWS, Governance & security.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 6 dremoved the moment Amazon closes it
