Amazon · amazon.jobs · checked today
Data Engineer III, Amazon Manufacturing Services (AMS)
Do you want to turn manufacturing data into decisions that move physical parts through a factory? Amazon Manufacturing Services (AMS) runs 135+ machines producing custom parts for over 100 Amazon organizations, and nearly every machine, order, and operator action generates data worth analyzing. You will join a small, growing data engineering team that owns the pipelines, warehouse, dashboards, and ML workflows that turn raw signals from our services and enterprise systems into throughput, utilization, and quality insights for shop floor users and AMS leadership.
Skills, with evidence
- Data modelling
Experience with data modeling, warehousing and building ETL pipelines
must have - Python
Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
must have - SQL
Experience with SQL
must have - Spark
Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
must have - Warehousing
Experience operating large data warehouses
must have - Streaming
- Design and operate data pipelines on AWS Glue (PySpark), Kinesis, S3, and EventBridge to ingest DynamoDB streams and enterprise system data into the AMS data lake
- AWS
- Design and operate data pipelines on AWS Glue (PySpark), Kinesis, S3, and EventBridge to ingest DynamoDB streams and enterprise system data into the AMS data lake
not practised here - Dashboards & BI
You will join a small, growing data engineering team that owns the pipelines, warehouse, dashboards, and ML workflows that turn raw signals from our services and enterprise systems into throughput, utilization, and quality insights for shop floor users and AMS leadership.
not practised here - Governance & security
- Build ingestion and modeling layers for enterprise data sources including SAP S/4HANA, JobBoss, Siemens Teamcenter, and Dot Compliance
not practised here - Data quality
- Own data quality, lineage, and documentation across the AMS analytics stack
- Java
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
not practised here - Machine learning
- Build and deploy ML models and pipelines for manufacturing use cases such as demand forecasting, machine health prediction, and scheduling optimization
not practised here - Scala
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
SQL · Data quality
- Median delivery time per cityIntermediate
- Bucket deliveries into quartilesIntermediate
- Median order value without a median functionIntermediate
- New and repeat orders by monthIntermediate
- Every order against its customer's averageIntermediate
- Python: the data-wrangling round≈ 3 h
Python
- Diff two snapshots of a tableIntermediate
- Explode an array column into rowsIntermediate
- Flatten nested event payloadsIntermediate
- Pivot a long metrics table to wideIntermediate
- Choose what an incremental run should readIntermediate
- Data modelling: the round most people fail≈ 3 h
Data modelling · Warehousing
- Addresses that stay true to the pastIntermediate
- Seat holds and the release-night raceIntermediate
- Subscription warehouse grainIntermediate
- Campaign efficiencyIntermediate
- Catalogue: products, variants and sellersIntermediate
- Pipeline design: safe to run twice≈ 4 h
Streaming
- Parcel tracking pipelineIntermediate
- Five minutes behind the sourceIntermediate
- The source will not let youIntermediate
- The fact arrived firstIntermediate
- Ten terabytes of clicks, and a question about sessionsAdvanced
- Spark: read the plan Spark actually ran≈ 2 h
Spark
- broadcast() with auto-broadcast off, and the case where Spark ignores itIntermediate
- autoBroadcastJoinThreshold compares an estimate: flip a join with select() and one settingIntermediate
- Does Spark really run your EXISTS subquery once per row?Intermediate
- left_semi and left_anti: "customers who did / never did" without a full joinIntermediate
- A self-join on a real key that still multiplies rowsIntermediate
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Dashboards & BI, Governance & security, Java, Machine learning, Scala.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 8 dremoved the moment Amazon closes it
