

Schema-on-Read Without Regrets: Lessons Learned from Apache Spark Variant Shredding in Production
The Variant type in Apache Spark gets a bad rap as "a way to create bad data." In this session, Scott Haines and Geir Alstad make the opposite case with a real, high-stakes pipeline: using variants to extract schema-less XML data in relational databases and fanning out as analytical tables.
We'll show how Variant shredding infers schema at runtime, shreds deeply nested structures into fast, indexable streaming tables, and delivers a clean Kimball star schema on the other side. This webinar will teach you honest operational lessons on schema drift, graceful reloads, source-system constraints, and the "Medallion Spoke" architecture that keeps data streaming from source to gold.
Speaker:
Geir Alstad is an experienced data warehouse developer and architect with 20 years of experience from financial services, public transport analysis, statistics and data science, business development. His current role has him responsible for building and architecting a large scale data lakehouse integration built on Databricks at Gabler AS. Over the last 3.5 years, he has created solutions using Lakehouse Well Architected Framework and oversees the implementation of enterprise wide adoption of the Databricks AI and Data platform. He is involved in large scale data migration projects in Gabler and oversees the transformational cloud adoption journey his company is on.
Host:
Scott Haines is a seasoned software engineer specializing in massive distributed data systems and streaming technologies. Over the past 18 years, he’s built and scaled data infrastructure at leading companies, including Yahoo!, Twilio, Nike, and, currently stepped into a role as a Staff Developer Advocate at Databricks.
Scott was a Databricks Beacon/MVP for five years before joining Databricks, he’s the author of books on Apache Spark and Delta Lake, and makes it his mission to help organizations successfully adopt open-source.