

SAIN Amsterdam Discussion Group: Astra & Chain-of-thought-monitorability
🤖 AI Safety Discussion Group
📍 Location: Oerknal, Science Park
🕠 Time: 5:30 PM
Discussion topic: Astra and chain-of-thought-monitorability
Some leaks before GPT-6-Astra's release suggested it might be less monitorable because of architectural changes. Now that it's out, there are a couple pieces of quick work suggesting this is indeed the case. We've picked out two of these that look particularly interesting
Astra can do a concerning amount with no chain of thought: https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-do-a-concerning-amount-with-no-chain-of-thought
Astra is much better at reasoning with filler tokens than previous models: https://www.lesswrong.com/posts/uvhuZHFtrgk8kNiZc/astra-is-much-better-at-reasoning-with-filler-tokens-than
Everyone is welcome, whether you are deeply involved in AI safety or simply curious about the topic!