Google · google.com · checked today
Data Engineer, Data Architecture and Engineering, gTech Strategy and Operations
As a Data Engineer in gData, you’ll design and build the foundational data infrastructure that powers the GBO data needs. You’ll manage complex, large-scale issues using Google’s proprietary tech stack to deliver a highly reliable Single Source of Truth. In this role, you will have the opportunity to architect innovative data pipelines and solutions that directly enable predictive problem-solving and AI-driven business decisions at scale.Google creates products and services that make the world a better place, and gTech’s role is to help bring them to life.
Skills, with evidence
- Python
Design, develop, test, and maintain reliable and scalable data pipelines and Extract, Transform, Load and Extract, Load, Transform (ETL/ELT) architectures using Google's distributed data systems (e.g., advanced SQL, Python).
must have - SQL
Design, develop, test, and maintain reliable and scalable data pipelines and Extract, Transform, Load and Extract, Load, Transform (ETL/ELT) architectures using Google's distributed data systems (e.g., advanced SQL, Python).
must have - Data modelling
Contribute to the modernization of the Google Ads Data Infrastructure (GDI) and Customer Data Platform (CDP), optimizing data models to ensure our Single Source of Truth remains robust and performant.
must have - Spark
1 year of experience with data processing software (e.g., Hadoop, Spark, Pig, Hive) and algorithms (e.g., MapReduce, Flume).
must have - Data quality
Advocate data quality by authoring clear technical design documents, executing code reviews, and proactively resolving complex bugs and support escalations.
must have - Java
Experience with database administration techniques or data engineering, as well as writing software in Java, C++, Python, Go, or JavaScript.
must have · not practised here - Warehousing
Experience in technical consulting and working with data warehouses, including data warehouse technical architectures, infrastructure components, ETL/ELT, and reporting/analytic tools and environments.
- Machine learning
Experience working with Big Data, information retrieval, data mining, or machine learning.
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
SQL · Data quality
- Airbnb: converting on a day with no published rateAdvanced
- Amazon: net revenue across two fan-outsAdvanced
- Build a percentile tableAdvanced
- Duplicate snapshot keysAdvanced
- Flag outlier ordersAdvanced
- Python: the data-wrangling round≈ 4 h
Python
- Data modelling: the round most people fail≈ 4 h
Data modelling · Warehousing
- Spark: read the plan Spark actually ran≈ 2 h
Spark
- applyInPandas per country: what Spark ships to Python, and the native rewriteAdvanced
- Filters you wrote in the wrong place: where Catalyst moves themAdvanced
- snappy, gzip or zstd: measure the Parquet codec trade-off yourselfAdvanced
- A correlated COUNT(*) subquery: the join Spark runs, and the count bug it avoidsAdvanced
- Find the four suspects in a slow pipeline (one of them is innocent)Advanced
- Say it out loud≈ 1 h
Not covered by the plan: Java, Machine learning.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves tomorrowremoved the moment Google closes it
