Cover Image for Do models know when they are being evaluated?
Cover Image for Do models know when they are being evaluated?
Avatar for Trajectory Labs
Presented by
Trajectory Labs
Hosted By
1 Went

Do models know when they are being evaluated?

Registration
Past Event
Welcome! To join the event, please register below.
About Event

Today's Topic

This week Giles Edkins will present his research on whether language models know when they are being evaluated.

One of the concerns with evaluating the alignment and capabilities of models is that they will detect that they are in an evaluation setting and modify their behaviour to appear less dangerous. At present little is known about what distinguishes real user conversations from automated evaluations and whether the language model is wise to these differences.

Giles has been researching what we can learn about this (together with Joe Needham and Govind Pimpale, mentored by Marius Hobbhahn as part of the MATS program).

Read more on Less Wrong here.

We welcome a variety of backgrounds, opinions and experience levels.

Event Schedule
6:00 to 6:45 - Networking and refreshments
6:45 to 8:00 - Main Presentation

Is there a topic you'd love to see us cover at a future event? Submit your suggestion here.

Location
30 Adelaide St E
Toronto, ON M5C 3G8, Canada
Enter the main lobby of the building and let the security staff know you are here for the AI meetup. You may need to show your RSVP on your phone. You will be directed to the 12th floor where the meetup is held. If you have trouble getting in, give Smitty a call at 647-424-4111.
Avatar for Trajectory Labs
Presented by
Trajectory Labs
Hosted By
1 Went