

San Diego Data Foundations Meetup
San Diego Data Foundations Meetup
AI is only as good as the data under it. San Diego Data Foundations brings together data engineers, architects and platform leaders to talk about what it takes to build data platforms that are ready for AI.
Hosted by Intuit, with speakers and practitioners from companies across the region, the evening covers:
Modern data platforms at scale: running distributed databases like Cassandra in production, with best practices for performance, reliability and day-to-day operations
The semantic layer: shared metrics, ontologies and knowledge graphs that give people, dashboards and AI agents the same business context
AI-ready data: quality, governance, metadata and context that turn raw data into something LLMs and agents can use
Program
Knowledge Graphs: A Shared Model of Meaning for Connected Data
Harsha Srinivasan, Staff Engineer, Intuit
Your systems already describe the same customer, but each one names it differently (customer_id, ClientID, owner) and nothing records that they match. This talk shows how to fix that with an ontology.
You'll learn how to derive a first-draft ontology from schemas you already own. You'll also see why a native graph database like Neo4j is a more practical way to run it in production than RDF/SPARQL. The result is a knowledge graph that AI agents can read directly to find data and ground their answers.
Single-Digit p99s: How Java 21, Direct I/O, and Cursor Compaction Change Cassandra Performance and Cost
Jon Haddad, Apache Cassandra Committer & Independent Consultant
Cassandra's tail latency has historically been dominated by a few predictable sources: garbage collection pauses, page cache poisoning, and compaction competing with reads for CPU and memory. Recent work in Cassandra goes after each of these directly, and together the changes bring p99 latencies down into the low single-digit milliseconds.
Java 21 brings Generational ZGC, which keeps pause times short while cutting the CPU and memory overhead of earlier ZGC versions.
Direct I/O lets Cassandra bypass the page cache, so compaction and streaming no longer evict the hot data your reads depend on.
Cursor-based compaction reduces object allocation during compaction, which lowers GC pressure and frees CPU for serving requests.
Jon explains how each change works, how to measure its effect on your own workload, and what configuration you need to get the benefit. He then connects the latency improvements to cost: when nodes spend less time on GC and compaction, you can handle the same traffic with fewer or smaller instances. You'll see how to estimate those savings and where the tradeoffs are.
AI for your Cassandra operations: something to embrace?
Johnny Miller, Co-founder & Chief Architect, AxonOps
Giving an AI access to a production Cassandra cluster sounds like the opening line of a post-mortem. It can also get from a 3am alert to a probable root cause while the on-call engineer is still looking for their laptop.
Johnny will give a live demo of what AI does for cluster operations today, what's around the corner, and where it should make an operator nervous. Embrace it, be cautious, or both? Bring your scepticism and your questions.
More speakers to be announced.
Expect real-world talks, honest lessons learned, and good conversation with peers solving the same problems. Food and drinks provided.