Google · google.com · checked today
Senior Data Engineer, gTech Risk
Google creates products and services that make the world a better place, and gTech’s role is to help bring them to life. Our teams of trusted advisors support customers globally. Our solutions are rooted in our technical skill, product expertise, and a thorough understanding of our customers’ complex needs.
Skills, with evidence
- Data modelling
5 years of experience designing data pipelines, and dimensional data modeling for synch and asynch system integration and implementation using internal (e.g., Flume, etc.) and external stacks (DataFlow, Spark, etc.).
must have - Python
Experience with Python or R for statistical analysis and model development.
must have - SQL
Experience in SQL and data visualization tools (e.g., Tableau, Looker).
must have - Spark
5 years of experience designing data pipelines, and dimensional data modeling for synch and asynch system integration and implementation using internal (e.g., Flume, etc.) and external stacks (DataFlow, Spark, etc.).
must have - Dashboards & BI
Build and manage executive dashboards for business reviews, pipeline management, and enterprise risk posture reporting.
not practised here - Streaming
5 years of experience designing data pipelines, and dimensional data modeling for synch and asynch system integration and implementation using internal (e.g., Flume, etc.) and external stacks (DataFlow, Spark, etc.).
- Failure handling
Partner with core Data Science and Engineering teams to transition our risk analysis from reactive (incident based) to predictive (threat based).
- Governance & security
Experience working within an Enterprise Risk Management (ERM) or Governance, Risk, and Compliance (GRC) framework.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
SQL
- Median delivery time per cityIntermediate
- Bucket deliveries into quartilesIntermediate
- Median order value without a median functionIntermediate
- New and repeat orders by monthIntermediate
- Every order against its customer's averageIntermediate
- Python: the data-wrangling round≈ 3 h
Python
- Diff two snapshots of a tableIntermediate
- Explode an array column into rowsIntermediate
- Flatten nested event payloadsIntermediate
- Pivot a long metrics table to wideIntermediate
- Choose what an incremental run should readIntermediate
- Data modelling: the round most people fail≈ 3 h
Data modelling
- Addresses that stay true to the pastIntermediate
- Seat holds and the release-night raceIntermediate
- Subscription warehouse grainIntermediate
- Campaign efficiencyIntermediate
- Catalogue: products, variants and sellersIntermediate
- Pipeline design: safe to run twice≈ 4 h
Streaming · Failure handling
- The source will not let youIntermediate
- Parcel tracking pipelineIntermediate
- Is this change safe?Intermediate
- Five minutes behind the sourceIntermediate
- Changing a pipeline that’s already runningIntermediate
- Spark: read the plan Spark actually ran≈ 2 h
Spark
- broadcast() with auto-broadcast off, and the case where Spark ignores itIntermediate
- autoBroadcastJoinThreshold compares an estimate: flip a join with select() and one settingIntermediate
- Does Spark really run your EXISTS subquery once per row?Intermediate
- left_semi and left_anti: "customers who did / never did" without a full joinIntermediate
- A self-join on a real key that still multiplies rowsIntermediate
- Say it out loud≈ 1 h
25 drills · Intermediate + Advanced≈ 14 hours
Not covered by the plan: Dashboards & BI, Governance & security.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 13 dremoved the moment Google closes it
