DUSTIN VANNOY

Azure Data Platform Overview slides

By Dustin Vannoy Aug 18, 2023 / Leave a comment

I had the privilege to present for Creating Coding Careers, a great organization in the San Diego area that helps people get established in tech careers via apprenticeships and other programs. Above are the slides used in that presentation. Recommended Resources to learn Azure Data Platform Databricks Training https://www.databricks.com/learn Microsoft Learn Training https://learn.microsoft.com/en-us/training/paths/data-engineer-azure-databricks/ https://learn.microsoft.com/en-us/training/paths/get-started-data-engineering/ https://learn.microsoft.com/en-us/training/paths/get-started-fabric/… Continue Reading

Data + AI Summit 2023 – Data Engineer key takeaways

By Dustin Vannoy Jun 30, 2023 / Leave a comment

Data + AI Summit 2023 has just completed with many announcements and deep dives. I attended virtually this year but was just as excited as the in-person attendees for some of the new capabilities that were shared. After watching the keynote presentations and tracking additional posts about new features, I want to summarize the top… Continue Reading

Apache Spark DataKickstart: Read and Write with PySpark

By Dustin Vannoy Jun 21, 2023 / 1 Comment

Every Spark pipeline involves reading data from a data source or table. For data engineers we usually end the pipelines by writing the transformed data. In this tutorial we walk through some of the most common format and cloud storage locations for reading and writing with Spark. We’ll save some of the advanced Delta Lake… Continue Reading

Apache Spark DataKickstart: First Spark SQL Application

By Dustin Vannoy May 18, 2023 / 1 Comment

Get hands on with Spark SQL (no Python or Scala) to build your first data pipeline. In this video I walk you through how to read, transform, and write the NYC Taxi dataset with Spark SQL. This dataset can be found on Databricks, Azure Synapse, or downloaded from the web to wherever you run Apache… Continue Reading

Apache Spark DataKickstart: First PySpark Application

By Dustin Vannoy May 1, 2023 / 1 Comment

Get hands on with Python and PySpark to build your first data pipeline. In this video I walk you through how to read, transform, and write the NYC Taxi dataset which can be found on Databricks, Azure Synapse, or downloaded from the web to wherever you run Apache Spark. Once you have watched and followed… Continue Reading

Apache Spark DataKickstart – Introduction

By Dustin Vannoy May 1, 2023 / Leave a comment

In this video I provide introduction to Apache Spark as part of my YouTube course Apache Spark DataKickstart. This video covers why Spark is popular, what it really is, and a bit about ways to run Apache Spark. Please check out other videos in this series by selecting the relevant playlist or subscribe and turn… Continue Reading

Data Engineer Question and Answer

By Dustin Vannoy Jan 19, 2023 / 2 Comments

An aspiring data engineer recently reached out to me for some guidance on pivoting into the field from a software development background. The questions they asked are similar to what others have asked me in the past, so I decided to capture my responses here. I link to prior posts and other resources when possible… Continue Reading

Questions Answered: Parallel Load in Spark Notebook

By Dustin Vannoy Jan 9, 2023 / 1 Comment

I received many questions on my tutorial Ingest tables in parallel with an Apache Spark notebook using multithreading. In this video and post I address some of the questions that I couldn’t just answer in the YouTube comments. Watch the video for more complete answers but here are quick responses with links to examples where… Continue Reading

Getting Started with Spark Structured Streaming – Current 22

By Dustin Vannoy Oct 5, 2022 / Leave a comment

I am honored to speak at Current 22. The example notebook that I walk through towards the end is available at https://github.com/datakickstart/datakickstart-databricks-workspace/blob/main/stackoverflow/stackoverflow_streaming.py.

Snowflake on Azure – Load with Synapse Pipeline

By Dustin Vannoy Jun 30, 2022 / Leave a comment

If you choose to use Snowflake along with Azure for your data platform, you will have to make choices on how to load the data. Landing processed data into your data lake on Azure Data Lake Storage Gen2 (ADLS) is the first step that I recommend in most environments. I like this pattern because then… Continue Reading