

Accuracy Isn’t Everything: Good Models are Lurking Everywhere
Please join Charlottesville Data Science for the talk Accuracy Isn’t Everything: Good Models are Lurking Everywhere from Evzenie Coupkova, a mathematics PhD from Purdue. In this talk, instead of focusing on one “best” model, Evzenie will explore the broad landscape of possible models, consider what it means for a machine learning model to be “good” for a given problem, and examine the trade-offs between predictive accuracy and other desirable characteristics, like generalization and interpretability.
We'll be gathering in person at the UVA School of Data Science on the evening of Thursday, October 8. We look forward to seeing you there!
→ Are you signed up for the Charlottesville Data Science Substack? If not, please consider subscribing at https://news.cvilleds.org. It's free, and it's the primary way we share announcements and event updates with the data science, AI, and machine learning community in Charlottesville and Central Virginia. ←
About the talk
Data scientists often focus on a single "best" model — the one with the highest accuracy. But what if we looked at all possible models? Would many others achieve similar accuracy and how does this depend on the dataset?
This talk sits at the intersection of Cynthia Rudin's work on interpretability and Sanjeev Arora's research on optimization and generalization of neural networks.
We start with simple, intuitive cases: linear models applied to Gaussian mixtures. We explore where the high- and low-accuracy models are located in parameter space, and how the proportion of "good" models change as the distance between class means grows.
We then turn to randomly initialized neural networks, asking the same question: what fraction of them are "good" classifiers? Using Sanjeev Arora's matrix-based approach to measuring dataset complexity, we examine how that complexity affects the fraction of "good" networks. The results of experiments on real-life datasets with varying sample sizes and dimensionality will be presented.
Underlying both parts of the talk is a general principle that holds for any model-data pair: if you know "good" models are abundant, you can improve generalization without giving up much accuracy. You can also choose a model that better fits your needs, one that is easily interpretable and suitable for high-stakes decision environments (such as healthcare or the judicial system), one that's sparse, or one that satisfies constraints imposed by a human expert.
If you want to shift your perspective and explore the full landscape of models, this talk is for you.
About the speaker
Evzenie Coupkova holds a PhD in Mathematics from Purdue University and specializes in statistical learning theory, machine learning, and high-dimensional data analysis. Her research explores the dimensionality reduction, proportion of high-accuracy models and generalizability in connection with randomness and dataset labels. She has collaborated with industry partners and presented her work at conferences including IPAM-UCLA and SIAM MDS, blending mathematical rigor with practical data science applications.