Summer Sale Special Limited Time 65% Discount Offer - Ends in 0d 00h 00m 00s - Coupon code: ac4s65

A data engineer is cleaning a Bronze table that receives the same customer records from...

A data engineer is cleaning a Bronze table that receives the same customer records from multiple source systems. Duplicate rows have the same customer_id and email values but different ingestion_timestamp values. The Silver table should contain only one record for each unique combination of customer_id and email.

Which PySpark operation correctly deduplicates the records based on the business keys?

A.

df.groupBy( " customer_id " , " email " ).agg(max( " ingestion_timestamp " ).alias( " latest_ts " ))

B.

df.distinct()

C.

df.select( " customer_id " , " email " ).distinct()

D.

df.dropDuplicates([ " customer_id " , " email " ])

Databricks-Certified-Data-Engineer-Associate PDF/Engine
  • Printable Format
  • Value of Money
  • 100% Pass Assurance
  • Verified Answers
  • Researched by Industry Experts
  • Based on Real Exams Scenarios
  • 100% Real Questions
buy now Databricks-Certified-Data-Engineer-Associate pdf
Get 65% Discount on All Products, Use Coupon: "ac4s65"