Google · google.com · checked today
Data Engineer, YouTube Business Organization
gTech’s Product and Tools Operations team (gPTO) leverages deep user, operational, and technical insights to innovate Google's Ads products into customer experiences that are so intuitive (or automated) that they require no support at all. gPTO partners closely with gTech’s Support, Professional Services, Product Management, and Engineering teams to innovate and simplify our Ads products and build the productivity tools ecosystem for gTech users. With over 2 billion monthly logged-in users, YouTube has grown into a global community where people all over the world access information, share vide
Skills, with evidence
- Data modelling
3 years of experience designing data pipelines, and dimensional data modeling for synch and asynch system integration and implementation using internal (e.g., Flume, etc.) and external stacks (DataFlow, Spark, etc.).
must have - SQL
We use SQL and YouTube’s Extract, Transform, Load (ETL) systems to produce useful datasets, establish best practices for data sets and reporting, and develop a breadth of expertise in various data domains.
must have - Spark
3 years of experience designing data pipelines, and dimensional data modeling for synch and asynch system integration and implementation using internal (e.g., Flume, etc.) and external stacks (DataFlow, Spark, etc.).
must have - Data quality
Build and maintain data platforms to enable data reliability, data integrity, and data governance, enabling accurate, consistent, and trustworthy data sets.
must have - AWS
Experience with data warehouses, large-scale distributed data platforms, data lakes, artificial intelligence and Gen AI data applications.
not practised here - Streaming
3 years of experience designing data pipelines, and dimensional data modeling for synch and asynch system integration and implementation using internal (e.g., Flume, etc.) and external stacks (DataFlow, Spark, etc.).
- Cost & performance
Design, build, and optimize the data architecture and extract, transform, and load (ETL) pipelines.
- Governance & security
Build and maintain data platforms to enable data reliability, data integrity, and data governance, enabling accurate, consistent, and trustworthy data sets.
not practised here - Machine learning
Work closely with analysts to productionize and scale value-creating capabilities, including data integrations and transformations, model features, and statistical and machine learning models.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 1 h
SQL · Data quality
- Average order value by countryFoundations
- Line items per orderFoundations
- Revenue by buyer countryFoundations
- Customers who never orderedFoundations
- Hiring funnel by roleFoundations
- Data modelling: the round most people fail≈ 2 h
Data modelling
- Marketplace core entitiesFoundations
- Cinema seat bookingFoundations
- City parking baysFoundations
- Dating app matchesFoundations
- Food delivery ordersFoundations
- Pipeline design: safe to run twice≈ 4 h
Streaming
- Which day does it belong to?Foundations
- Parcel tracking pipelineIntermediate
- Five minutes behind the sourceIntermediate
- The source will not let youIntermediate
- The fact arrived firstIntermediate
- Spark: read the plan Spark actually ran≈ 1 h
Spark · Cost & performance
- HAVING vs WHERE: where does a filter after groupBy actually run?Foundations
- countDistinct vs approx_count_distinct: what the extra shuffle buysFoundations
- COUNT(*) vs COUNT(column): the null trapFoundations
- Grouping by two columns: what changes in the shuffle?Foundations
- Does the join type change the join strategy?Foundations
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Governance & security, Machine learning.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 13 dremoved the moment Google closes it
