Adobe · adobe.wd5.myworkdayjobs.com · checked today
Senior Data Engineer
The world of design is changing rapidly and the Pro Design team is leading that transformation. We are the Adobe organization behind Illustrator, InDesign, and emerging experiences that connect creativity, collaboration, and AI. Our teams are reimagining what professional design looks like for the next decade - building intelligent, connected tools that empower creators and teams to move faster without sacrificing craft.
Skills, with evidence
- Airflow / orchestration
Proficiency in SQL and Python, PySpark, Airflow, Databricks workflows.
must have - Data quality
Set and evolve governance standards regarding data quality, privacy, security, lineage, and SLAs in Unity Catalog.
must have - Python
Proficiency in SQL and Python, PySpark, Airflow, Databricks workflows.
must have - SQL
Proficiency in SQL and Python, PySpark, Airflow, Databricks workflows.
must have - Cost & performance
Build, scale and optimize the data architecture and ETL/pipeline on Databricks across Pro Design products, spanning shared platform data and in-app usage telemetry.
- Governance & security
Set and evolve governance standards regarding data quality, privacy, security, lineage, and SLAs in Unity Catalog.
not practised here - Spark
Build, scale and optimize the data architecture and ETL/pipeline on Databricks across Pro Design products, spanning shared platform data and in-app usage telemetry.
- Data modelling
A track record of partnering with Product, Engineering, and Data Science to turn ambiguous questions into reliable data models.
- Warehousing
Hands-on experience with Databricks and Unity Catalog (or comparable Lakehouse and governance tooling).
- AWS
Experience deploying and managing services and applications on Azure and Azure Cloud platforms.
not practised here - Azure
Experience deploying and managing services and applications on Azure and Azure Cloud platforms.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
Data quality · SQL
- Median delivery time per cityIntermediate
- Bucket deliveries into quartilesIntermediate
- Median order value without a median functionIntermediate
- New and repeat orders by monthIntermediate
- Every order against its customer's averageIntermediate
- Python: the data-wrangling round≈ 3 h
Python
- Diff two snapshots of a tableIntermediate
- Explode an array column into rowsIntermediate
- Flatten nested event payloadsIntermediate
- Pivot a long metrics table to wideIntermediate
- Choose what an incremental run should readIntermediate
- Data modelling: the round most people fail≈ 3 h
Data modelling · Warehousing
- Addresses that stay true to the pastIntermediate
- Seat holds and the release-night raceIntermediate
- Subscription warehouse grainIntermediate
- Campaign efficiencyIntermediate
- Catalogue: products, variants and sellersIntermediate
- Pipeline design: safe to run twice≈ 4 h
Airflow / orchestration
- Is this change safe?Intermediate
- Marketplace transactions at scaleIntermediate
- The source will not let youIntermediate
- Changing a pipeline that’s already runningIntermediate
- SLA-aware alerting flowIntermediate
- Spark: read the plan Spark actually ran≈ 2 h
Cost & performance · Spark
- broadcast() with auto-broadcast off, and the case where Spark ignores itIntermediate
- autoBroadcastJoinThreshold compares an estimate: flip a join with select() and one settingIntermediate
- Does Spark really run your EXISTS subquery once per row?Intermediate
- left_semi and left_anti: "customers who did / never did" without a full joinIntermediate
- A self-join on a real key that still multiplies rowsIntermediate
- Say it out loud≈ 1 h
Not covered by the plan: Governance & security, AWS, Azure.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 8 dremoved the moment Adobe closes it
