

DET Webinar: Stop Optimizing the Wrong Part of Your Spark Job
High utilization doesn’t necessarily mean an efficient Spark workload – and the layer where waste shows up is often not where you need to fix it. In this webinar, we’ll walk through a practical framework for investigating and optimizing Spark workloads across infrastructure, runtime, data access, and query plans, using real production cases to show the signals, root causes, and optimizations at each layer. We’ll also discuss what it takes to collect the runtime signals needed to apply this framework effectively, and the trade-offs between different approaches to monitoring Spark workloads.
Speakers
Ohad Raviv: CTO & Co-Founder, definity
Ohad is CTO and co-founder of definity, where he works on Spark optimization and agentic data engineering. Previously, he led the big data technology roadmap for PayPal’s global data science group, pioneering work across Hadoop and Spark observability, distributed graph infrastructure, and entity resolution. Ohad is an Apache Spark contributor.
Roy Daniel: CEO & Co-Founder, definity
Roy is CEO and co-founder of definity, an Agentic Data Engineering platform for the Lakehouse and Spark ecosystem. He works with enterprise data engineering teams to optimize cost, improve reliability, and accelerate developer velocity. Previously, he held data and product leadership roles across Fortune 500 companies, most recently at FIS.
Sponsor
Thank you to definity for sponsoring and supporting the community.
📚 About Data Engineer Things
Data Engineer Things (DET) is a global community built by data engineers for data engineers. Subscribe to the newsletter and follow us on LinkedIn to gain access to exclusive learning resources and networking opportunities, including articles, webinars, meetups, conferences, mentorship, and more.