Deep Spark internals (OOM, optimized joins, skew + salting), dimensional modeling, and batch+streaming (Spark+Kafka) design.
Cloud / platform
AWS; Spark, Hadoop, Hive, Kafka
Reported difficulty
Not reported
Assessment (Python+SQL) → HM → onsite (4 rounds + code pairing). Heavy Spark: OOM fixes, optimize a Spark app, optimized joins, data skew + salting; DSA (missing-number-style); dimensional modeling; LLD (OOP/classes).
JoinTaro (Jun 2025) confirming the multi-round loop.
Answered questions on what Expedia actually asks — each with the short answer, the answer that gets you rejected, and the follow-up.
Then build one and have it reviewed: design a daily orders pipeline or model a marketplace. No account.