

How do we solve alignment?
How Do We Solve Alignment?
This Tuesday, Edward Y. (Geodesic Research, formally UK AISI), will give an informal talk on "Will the character of AI Assistants become more incoherent over time?"
From the man himself:
The talk will discuss the possibility that AI assistants will become less globally coherent in their characters and behaviours as post-training compute is scaled up.
I will cover recent results and observations that point in the direction of AI assistants becoming increasingly incoherent in their characters over time, and discuss various explanations for why this might be the case.
I will then discuss various standard intuitions for why we would expect AI systems to be coherent, and whether those intuitions are valid.
Finally, I will note open-problems in this area and possibilities for empirical research.