Amazon · amazon.jobs · checked today
Data Engineer, Amazon Music, Amazon Music Finance
Amazon Music is an immersive audio entertainment service that deepens connections between fans, artists, and creators. From personalized music playlists to exclusive podcasts, concert livestreams to artist merch, Amazon Music is innovating at some of the most exciting intersections of music and culture. We offer experiences that serve all listeners with our different tiers of service: Prime members get access to all the music in shuffle mode, and top ad-free podcasts, included with their membership;
Skills, with evidence
- Airflow / orchestration
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
must have - Data modelling
Develop data models to support business intelligence, delivering actionable insights and interactive reports to end-users.
must have - Python
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
must have - SQL
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
must have - Warehousing
Devise and implement an efficient, scalable data warehousing solution on AWS, utilizing appropriate NoSQL/SQL storage and database technologies for both structured and unstructured data.
must have - AWS
Engage in collaborative efforts with cross-functional teams — data scientists, business intelligence engineers, and Finance Managers — to architect a state-of-the-art data analytics platform on AWS using the AWS Cloud Development Kit (CDK).
must have · not practised here - Machine learning
Enable advanced analytics, machine learning, and generative AI capabilities within the platform, extracting predictive and prescriptive insights through tools like EMR and SageMaker.
must have · not practised here - Spark
Enable advanced analytics, machine learning, and generative AI capabilities within the platform, extracting predictive and prescriptive insights through tools like EMR and SageMaker.
- Cost & performance
Continuously monitor and optimize the performance of data pipelines, databases, and applications, ensuring low-latency data access for analytics and machine learning tasks.
- Data quality
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
- Failure handling
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
- Streaming
- Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
- Governance & security
Implement robust security measures and ensure data compliance with internal requirements, industry standards, and regulations to safeguard sensitive information.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 1 h
SQL · Data quality
- Average order value by countryFoundations
- Line items per orderFoundations
- Revenue by buyer countryFoundations
- Customers who never orderedFoundations
- Hiring funnel by roleFoundations
- Python: the data-wrangling round≈ 2 h
Python
- Clean spreadsheet headers into column namesFoundations
- Events normalization jobFoundations
- List the distinct composite keysFoundations
- Deduplicate rows, first one winsFoundations
- Render a byte count for humansFoundations
- Data modelling: the round most people fail≈ 2 h
Data modelling · Warehousing
- Marketplace core entitiesFoundations
- Cinema seat bookingFoundations
- City parking baysFoundations
- Dating app matchesFoundations
- Food delivery ordersFoundations
- Pipeline design: safe to run twice≈ 3 h
Airflow / orchestration · Failure handling · Streaming
- After the first one finishesFoundations
- Run it for last TuesdayFoundations
- The history nobody keptFoundations
- The spreadsheet is a dependencyFoundations
- Too slow by morningFoundations
- Spark: read the plan Spark actually ran≈ 1 h
Spark · Cost & performance
- HAVING vs WHERE: where does a filter after groupBy actually run?Foundations
- countDistinct vs approx_count_distinct: what the extra shuffle buysFoundations
- COUNT(*) vs COUNT(column): the null trapFoundations
- Grouping by two columns: what changes in the shuffle?Foundations
- Does the join type change the join strategy?Foundations
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Machine learning, Governance & security.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 12 dremoved the moment Amazon closes it
