Deep Spark internals and cluster sizing, scenario SQL (self-joins, window-with-range), and classic coding problems.
Cloud / platform
Spark, Hadoop, Hive, Airflow, Scala
Reported difficulty
Not reported
OA → technical (SQL: dedup, IPL self-join, window-with-range, bank-statement design; Spark: join optimization, SCD types, file formats, 5min→2hr perf debug, repartition vs coalesce, cluster sizing 500GB; coding: word freq, subarray sum K, merge K sorted lists) → behavioral.
Answered questions on what Paytm actually asks — each with the short answer, the answer that gets you rejected, and the follow-up.
Then build one and have it reviewed: design a daily orders pipeline or model a marketplace. No account.