Data Engineer · Mid · Amazon

the posting, the skills it asks for, and the plan to get there

Amazon · amazon.jobs · checked today

Data Engineer - Music DISCO, Music DISCO

Mexico City, Mexico City, MEXMid3+ yrsData Engineerposted 6 d ago

Amazon Music is awash in data! To help make sense of it all, the DISCO (Data, Insights, Science & Optimization) team: (i) enables the Consumer Product Tech org make data driven decisions that improve the customer retention, engagement and experience on Amazon Music. We build and maintain automated self-service data solutions, data science models and deep dive difficult questions that provide actionable insights.

Skills, with evidence

  • Airflow / orchestration
    Your team will manage the data exchange store (Data Lake) and EMR/Spark processing layer using Airflow as orchestrator.must have
  • Data modelling
    Duties include big data design and analysis, data modeling, and development, deployment, and operations of big data pipelines.must have
  • Python
    Develop robust and scalable data pipelines using SQL/PySpark/Airflow to efficiently ingest, process, and transform large volumes of data from various sources into a structured format, ensuring data quality and integrity.must have
  • SQL
    Develop robust and scalable data pipelines using SQL/PySpark/Airflow to efficiently ingest, process, and transform large volumes of data from various sources into a structured format, ensuring data quality and integrity.must have
  • Spark
    We collect billions of events a day, manage petabyte scale data on Redshift and S3, and develop data pipelines using Spark/Scala EMR, SQL based ETL, Airflow services.must have
  • Warehousing
    Design and implement an efficient and scalable data warehousing solution on AWS, using appropriate NoSQL/SQL storage and database technologies for structured and unstructured data.must have
  • AWS
    We deal in AWS technologies like Redshift, S3, EMR, EC2, DynamoDB, Kinesis Firehose, and Lambda.must have · not practised here
  • Machine learning
    Enable advanced analytics and machine learning capabilities within the platform to derive predictive and prescriptive insights from the data through tools like EMR/SageMaker Notebooks.must have · not practised here
  • Cost & performance
    We reduce the cost in time and effort of analysis, data set building, model building, and user segmentation.
  • Streaming
    We deal in AWS technologies like Redshift, S3, EMR, EC2, DynamoDB, Kinesis Firehose, and Lambda.
  • Data quality
    -Develop robust and scalable data pipelines using SQL/PySpark/Airflow to efficiently ingest, process, and transform large volumes of data from various sources into a structured format, ensuring data quality and integrity.
  • Governance & security
    -Implement robust security measures and ensure data compliance with internal requirements, industry standards, and regulations to safeguard sensitive information.not practised here

Your plan

  1. The SQL screen: correct, then fast

    SQL · Data quality

    1 h
  2. 2 h
  3. Data modelling: the round most people fail

    Data modelling · Warehousing

    2 h
  4. Pipeline design: safe to run twice

    Airflow / orchestration · Streaming

    3 h
  5. 1 h
  6. 1 h
25 drills · Foundations + Intermediate10 hours

Not covered by the plan: AWS, Machine learning, Governance & security.

Readiness

Counted from drills you have completed anywhere on D8LooP.

leaves in 7 dremoved the moment Amazon closes it

Data Engineer - Music DISCO, Music DISCO at Amazon: skills and prep plan · D8LooP