Nebius · careers.nebius.com · checked today
Senior Data Engineer
<div class="content-intro"><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p>
Skills, with evidence
- Airflow / orchestration
Orchestrate pipelines using a workflow orchestration framework (e.g., Airflow or equivalent).
must have - Cost & performance
Optimize pipelines for performance, reliability, and cost efficiency.
must have - Data quality
Implement data transformations, validation, and data quality checks.
must have - Idempotency & backfills
Develop stateless, idempotent pipelines that are resilient to retries, failures, and infrastructure interruptions.
must have - Python
Design, build, and own production-grade data pipelines using Python and SQL.
must have - SQL
Design, build, and own production-grade data pipelines using Python and SQL.
must have - Kubernetes
3+ years of experience running workloads on Kubernetes.
must have · not practised here - Machine learning
<p>This is a hands-on data engineering role, focused on designing, implementing, and maintaining reliable data flows for analytics and machine learning.
not practised here - Failure handling
<li>Develop stateless, idempotent pipelines that are resilient to retries, failures, and infrastructure interruptions.</li>
- Spark
<li>Experience building data pipelines using Apache Spark or similar distributed processing frameworks.</li>
- Terraform
<li>Use Infrastructure as Code only to provision and manage the infrastructure required to run pipelines.</li>
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
Data quality · SQL
- Median delivery time per cityIntermediate
- Bucket deliveries into quartilesIntermediate
- Median order value without a median functionIntermediate
- New and repeat orders by monthIntermediate
- Every order against its customer's averageIntermediate
- Python: the data-wrangling round≈ 3 h
Python
- Diff two snapshots of a tableIntermediate
- Explode an array column into rowsIntermediate
- Flatten nested event payloadsIntermediate
- Pivot a long metrics table to wideIntermediate
- Choose what an incremental run should readIntermediate
- Pipeline design: safe to run twice≈ 4 h
Airflow / orchestration · Idempotency & backfills · Failure handling
- Is this change safe?Intermediate
- Marketplace transactions at scaleIntermediate
- Five minutes behind the sourceIntermediate
- The source will not let youIntermediate
- Changing a pipeline that’s already runningIntermediate
- Spark: read the plan Spark actually ran≈ 2 h
Cost & performance · Spark
- broadcast() with auto-broadcast off, and the case where Spark ignores itIntermediate
- autoBroadcastJoinThreshold compares an estimate: flip a join with select() and one settingIntermediate
- Does Spark really run your EXISTS subquery once per row?Intermediate
- left_semi and left_anti: "customers who did / never did" without a full joinIntermediate
- A self-join on a real key that still multiplies rowsIntermediate
- Say it out loud≈ 1 h
Not covered by the plan: Kubernetes, Machine learning, Terraform.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 13 dremoved the moment Nebius closes it
