Verifying Science in the Age of AI
Science has a verification problem that predates AI, and AI is about to make it worse. Roughly a third of landmark psychology findings replicate. Over 10,000 papers were retracted in 2023, a record. Irreproducible preclinical research costs the US tens of billions annually. Now AI-generated submissions are beginning to flood a peer review system that runs on unpaid, overloaded human labor.
AI can be part of the solution to this problem. AI can now do the tedious work that made verification impossible to scale: extracting claims and statistics from papers, running forensic checks (statistical anomalies, image manipulation, deviations from preregistration), reproducing analyses where code and data exist, and updating evidence syntheses continuously as new results arrive. In a recent benchmark, an ensemble of AI reviewers caught 93 of 100 errors deliberately planted in published psychology papers, at a few cents per error caught. The components of a living evidence layer (an open, continuously updated map of scientific claims and how much to trust each one) are already being built by startups, academic labs, and nonprofits. What's missing is integration and evaluation.
This talk lays out what such a system looks like, what it changes for funders, agencies, journalists, and the public, and where government action matters most: turning public archives like PubMed into structured data, commissioning open benchmarks for claim extraction and error detection, and keeping the result a public good rather than a publisher-owned asset.
Biography:
Paul Litvak is the founder and Executive Director of the Robyn Dawes Institute and a Visiting Scholar at UC Berkeley. Over 15 years in industry he worked at Meta, Google, and Airbnb as a data scientist and then leading Machine Learning teams, and co-founded Rhythmic Health, a venture-backed biosensing startup. He has a PhD in Behavioral Decision Research from Carnegie Mellon. His current work involves building AI systems that can weigh the balance of scientific evidence the way a careful, critical scientist would—flagging numbers that don't add up and assessing methodology across entire bodies of research. This will allow the people and organizations making high-stakes health and philanthropic decisions to know what the evidence actually supports.
About Your Hosts:
The Robyn Dawes Institute for the Improvement of Science makes research quality transparent and usable at scale, enabling researchers, journalists, policymakers, and the public to assess and act on evidence. Its independent tools analyze papers, run rigorous checks, and surface the most reliable findings.
The DC Industrial Policy Salon convenes leading experts and policymakers to discuss issues of science, technology, and industrial production in Washington, DC.