Amazon · amazon.jobs · checked today
Data Engineer, Amazon Music, Amazon Music Finance
Skills, with evidence
- Airflow / orchestration
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
must have - Data modelling
Develop data models to support business intelligence, delivering actionable insights and interactive reports to end-users.
must have - Python
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
must have - SQL
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
must have - Warehousing
Devise and implement an efficient, scalable data warehousing solution on AWS, utilizing appropriate NoSQL/SQL storage and database technologies for both structured and unstructured data.
must have - AWS
Engage in collaborative efforts with cross-functional teams — data scientists, business intelligence engineers, and Finance Managers — to architect a state-of-the-art data analytics platform on AWS using the AWS Cloud Development Kit (CDK).
must have · not practised here - Machine learning
Enable advanced analytics, machine learning, and generative AI capabilities within the platform, extracting predictive and prescriptive insights through tools like EMR and SageMaker.
must have · not practised here - Spark
Enable advanced analytics, machine learning, and generative AI capabilities within the platform, extracting predictive and prescriptive insights through tools like EMR and SageMaker.
- Cost & performance
Continuously monitor and optimize the performance of data pipelines, databases, and applications, ensuring low-latency data access for analytics and machine learning tasks.
- Data quality
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
- Failure handling
Construct resilient and scalable data pipelines using SQL/PySpark/Airflow to ingest, process, and transform substantial data volumes from diverse sources into a structured format, ensuring data quality and integrity.
- Streaming
- Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
- Governance & security
Implement robust security measures and ensure data compliance with internal requirements, industry standards, and regulations to safeguard sensitive information.
not practised here
Your plan
- SQL
drills at the Mid level
- Python
drills at the Mid level
- Modelling
drills at the Mid level
- Pipelines
drills at the Mid level
- Spark
drills at the Mid level
The full plan, drill by drill, is on the job page.
leaves in 12 dremoved the moment Amazon closes it
