Okta · okta.com · checked today
Senior Data Engineer
<div class="content-intro"><p><strong>Secure Every Identity, from AI to Human<br><br></strong>Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes.
Skills, with evidence
- Airflow / orchestration
Expert-level knowledge of SQL, ETL tools such as Airflow and dbt, and relational/columnar MPP databases like Snowflake or Redshift.
must have - Data quality
Implement comprehensive data governance frameworks to ensure data lineage, compliance, and quality across all analytical platforms.
must have - SQL
Expert-level knowledge of SQL, ETL tools such as Airflow and dbt, and relational/columnar MPP databases like Snowflake or Redshift.
must have - dbt
Expert-level knowledge of SQL, ETL tools such as Airflow and dbt, and relational/columnar MPP databases like Snowflake or Redshift.
must have - AWS
Hands-on experience with AWS (S3, Lambda, EMR, EC2, EKS).
must have · not practised here - Kubernetes
Proficiency with Docker and Kubernetes for packaging and distributing applications.
must have · not practised here - Machine learning
Engineer robust data pipelines that fuel high-performance AI and ML models by delivering reliable, feature-rich datasets.
must have · not practised here - Terraform
Utilize Terraform to manage infrastructure as code, ensuring consistent and reproducible environments.
must have · not practised here - Spark
<p><span style="font-family: arial, helvetica, sans-serif;"><strong>Design and Develop Scalable Platforms:</strong> Build and maintain scalable data platforms using AWS, Snowflake, dbt, and Databricks.</span></p>
- Governance & security
<p><span style="font-family: arial, helvetica, sans-serif;"><strong>Data Governance &
not practised here - Cost & performance
<p><span style="font-family: arial, helvetica, sans-serif;">Familiarity with cloud cost optimization (FinOps)</span></p>
- Warehousing
<p><span style="font-family: arial, helvetica, sans-serif;"><strong>Modern Data Architecture:</strong> Experience with lakehouse architectures (Databricks) and file formats like Iceberg and Delta.</span></p>
Your plan
- The SQL screen: correct, then fast≈ 2 h
Data quality · SQL
- Median delivery time per cityIntermediate
- Bucket deliveries into quartilesIntermediate
- Median order value without a median functionIntermediate
- New and repeat orders by monthIntermediate
- Every order against its customer's averageIntermediate
- Data modelling: the round most people fail≈ 3 h
dbt · Warehousing
- Addresses that stay true to the pastIntermediate
- Subscription warehouse grainIntermediate
- Campaign efficiencyIntermediate
- Chats, members and read receiptsIntermediate
- Churn that survives an argumentIntermediate
- Pipeline design: safe to run twice≈ 4 h
Airflow / orchestration
- Is this change safe?Intermediate
- Marketplace transactions at scaleIntermediate
- The source will not let youIntermediate
- Changing a pipeline that’s already runningIntermediate
- SLA-aware alerting flowIntermediate
- Spark: read the plan Spark actually ran≈ 2 h
Spark · Cost & performance
- broadcast() with auto-broadcast off, and the case where Spark ignores itIntermediate
- autoBroadcastJoinThreshold compares an estimate: flip a join with select() and one settingIntermediate
- Does Spark really run your EXISTS subquery once per row?Intermediate
- left_semi and left_anti: "customers who did / never did" without a full joinIntermediate
- A self-join on a real key that still multiplies rowsIntermediate
- Say it out loud≈ 1 h
Not covered by the plan: AWS, Kubernetes, Machine learning, Terraform, Governance & security.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 2 dremoved the moment Okta closes it
