Salesforce · salesforce.wd12.myworkdayjobs.com · checked today
Architect, Data Platform — AgentExchange
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Salesforce is the #1 AI CRM, where humans with agents drive customer success together. And innovation isn’t a buzzword — it’s a way of life.
Skills, with evidence
- Data modelling
Deep architecture experience in at least three of : lakehouse / warehouse design, streaming + batch pipelines, dimensional and event modeling, feature stores, model serving.
must have - Streaming
Pipelines and contracts. Streaming and batch ingestion, schema governance, data contracts enforced across every AgentExchange engineering team, and pipeline reliability SLOs.
must have - Governance & security
Data security and governance as a first-class skill: PII classification, multi-tenant isolation, fine-grained access control, GDPR / CCPA, lineage and audit, and the security implications of LLM / agent access patterns.
must have · not practised here - Machine learning
ML platform. Feature store, training and serving infrastructure, evaluation, and monitoring.
must have · not practised here - Data quality
Partner-facing dashboards refresh on a documented SLO, and instrumentation completeness is measurable and enforced.
must have - AWS
Cloud-native data infrastructure: Snowflake, BigQuery, Redshift, or Databricks; AWS-based platforms.
must have · not practised here - Dashboards & BI
Own the data and intelligence architecture for AgentExchange — the marketplace where partners list, sell, and operate Agentforce, MuleSoft, Tableau, and Slack solutions.
not practised here - Failure handling
Feature store, training and serving infrastructure, evaluation, and monitoring.
- Cost & performance
LLM systems experience, in production: RAG, embeddings and vector stores, prompt and context engineering, offline and online evaluation, cost and latency tuning, hallucination and safety controls.
- SQL
Cloud-native data infrastructure: Snowflake, BigQuery, Redshift, or Databricks;
- Schema evolution & contracts
Streaming and batch ingestion, schema governance, data contracts enforced across every AgentExchange engineering team, and pipeline reliability SLOs.
- Spark
Cloud-native data infrastructure: Snowflake, BigQuery, Redshift, or Databricks;
- Warehousing
Deep architecture experience in at least three of : lakehouse / warehouse design, streaming + batch pipelines, dimensional and event modeling, feature stores, model serving.
- GCP
Cloud-native data infrastructure: Snowflake, BigQuery, Redshift, or Databricks;
not practised here
Your plan
- The SQL screen: correct, then fast≈ 2 h
Data quality · SQL
- Airbnb: converting on a day with no published rateAdvanced
- Amazon: net revenue across two fan-outsAdvanced
- Build a percentile tableAdvanced
- Duplicate snapshot keysAdvanced
- Flag outlier ordersAdvanced
- Data modelling: the round most people fail≈ 4 h
Data modelling · Warehousing
- Pipeline design: safe to run twice≈ 4 h
Streaming · Failure handling · Schema evolution & contracts
- The payload changed overnightAdvanced
- Fourteen months of wrong numbersAdvanced
- Ten terabytes of clicks, and a question about sessionsAdvanced
- Creator payout runAdvanced
- Delete me from everywhereAdvanced
- Spark: read the plan Spark actually ran≈ 2 h
Cost & performance · Spark
- applyInPandas per country: what Spark ships to Python, and the native rewriteAdvanced
- Filters you wrote in the wrong place: where Catalyst moves themAdvanced
- snappy, gzip or zstd: measure the Parquet codec trade-off yourselfAdvanced
- A correlated COUNT(*) subquery: the join Spark runs, and the count bug it avoidsAdvanced
- Find the four suspects in a slow pipeline (one of them is innocent)Advanced
- Say it out loud≈ 1 h
Not covered by the plan: Governance & security, Machine learning, AWS, Dashboards & BI, GCP.
Readiness
Counted from drills you have completed anywhere on D8LooP.
leaves in 8 dremoved the moment Salesforce closes it
