Tech Talk archive · Part 1 of 9
Start data science with a Databricks overview
An archived introductory Tech Talk covering the Databricks workflow and its place in the nine-part Data Science for Dummies series.
This introductory 2019 Tech Talk placed Databricks between raw data and a working machine-learning workflow. It was designed for readers who had heard of Spark but had not yet used a managed Spark workspace.
The original slide player is no longer embedded. Its Slideshare reference was 150116462.
What the session introduced
- Spark as a distributed processing engine.
- A shared workspace for notebooks and data work.
- The roles of Python, SQL, Scala, and R in the platform at that time.
- The handoff from data preparation to experiments and production work.
These points describe the 2019 session. Check current Databricks documentation before selecting a runtime, language, library, or deployment pattern.
The nine-part series
- Data Science overview with Databricks
- Titanic survival prediction with Azure Machine Learning Studio and Kaggle
- Data engineering with the Titanic dataset, Databricks, and Python
- Titanic with Databricks and Spark ML
- Titanic with Databricks and Azure Machine Learning
- Titanic with Databricks and AutoML
- Titanic with Databricks and MLflow
- Titanic with .NET and ML.NET
- Deployment, DevOps, MLOps, and production
Use the series archive to find the retained posts.