Amazon · amazon.jobs · checked today
Data Engineer - Music DISCO, Music DISCO
Amazon Music is awash in data! To help make sense of it all, the DISCO (Data, Insights, Science & Optimization) team: (i) enables the Consumer Product Tech org make data driven decisions that improve the customer retention, engagement and experience on Amazon Music. We build and maintain automated self-service data solutions, data science models and deep dive difficult questions that provide actionable insights.
Skills, with evidence
- Airflow / orchestration
Your team will manage the data exchange store (Data Lake) and EMR/Spark processing layer using Airflow as orchestrator.
must have - Data modelling
Duties include big data design and analysis, data modeling, and development, deployment, and operations of big data pipelines.
must have - Python
Develop robust and scalable data pipelines using SQL/PySpark/Airflow to efficiently ingest, process, and transform large volumes of data from various sources into a structured format, ensuring data quality and integrity.
must have - SQL
Develop robust and scalable data pipelines using SQL/PySpark/Airflow to efficiently ingest, process, and transform large volumes of data from various sources into a structured format, ensuring data quality and integrity.
must have - Spark
We collect billions of events a day, manage petabyte scale data on Redshift and S3, and develop data pipelines using Spark/Scala EMR, SQL based ETL, Airflow services.
must have - Warehousing
Design and implement an efficient and scalable data warehousing solution on AWS, using appropriate NoSQL/SQL storage and database technologies for structured and unstructured data.
must have - AWS
We deal in AWS technologies like Redshift, S3, EMR, EC2, DynamoDB, Kinesis Firehose, and Lambda.
must have · not practised here - Machine learning
Enable advanced analytics and machine learning capabilities within the platform to derive predictive and prescriptive insights from the data through tools like EMR/SageMaker Notebooks.
must have · not practised here - Cost & performance
We reduce the cost in time and effort of analysis, data set building, model building, and user segmentation.
- Streaming
We deal in AWS technologies like Redshift, S3, EMR, EC2, DynamoDB, Kinesis Firehose, and Lambda.
- Data quality
-Develop robust and scalable data pipelines using SQL/PySpark/Airflow to efficiently ingest, process, and transform large volumes of data from various sources into a structured format, ensuring data quality and integrity.
- Governance & security
-Implement robust security measures and ensure data compliance with internal requirements, industry standards, and regulations to safeguard sensitive information.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 1 h
SQL · Data quality
- Average order value by countryFoundations
- Line items per orderFoundations
- Revenue by buyer countryFoundations
- Customers who never orderedFoundations
- Hiring funnel by roleFoundations
- Python: the data-wrangling round≈ 2 h
Python
- Clean spreadsheet headers into column namesFoundations
- Events normalization jobFoundations
- List the distinct composite keysFoundations
- Deduplicate rows, first one winsFoundations
- Render a byte count for humansFoundations
- Data modelling: the round most people fail≈ 2 h
Data modelling · Warehousing
- Marketplace core entitiesFoundations
- Cinema seat bookingFoundations
- City parking baysFoundations
- Dating app matchesFoundations
- Food delivery ordersFoundations
- Pipeline design: safe to run twice≈ 3 h
Airflow / orchestration · Streaming
- After the first one finishesFoundations
- Run it for last TuesdayFoundations
- The history nobody keptFoundations
- The spreadsheet is a dependencyFoundations
- Too slow by morningFoundations
- Spark: read the plan Spark actually ran≈ 1 h
Spark · Cost & performance
- HAVING vs WHERE: where does a filter after groupBy actually run?Foundations
- countDistinct vs approx_count_distinct: what the extra shuffle buysFoundations
- COUNT(*) vs COUNT(column): the null trapFoundations
- Grouping by two columns: what changes in the shuffle?Foundations
- Does the join type change the join strategy?Foundations
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Machine learning, Governance & security.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 7 dremoved the moment Amazon closes it
