Published: May 5, 2025
1
4
29

1/6: A recent paper shows that that LLMs are "self aware": when trained to exhibit a behavior like "risk taking", LLMs self report being risky. In a recent blog post, we explore what's happening here: some self awareness behaviors are caused by a simple learned steering vector!🧵

2/6: We study models finetuned with LoRA to be risk taking or risk avoidant. We find that 1 layer of LoRA is enough; when we investigate this LoRA, it turns out to just add a steering vector! The safety steering vector even has high cosine sim to "safety" unembedding tokens.

Image in tweet by Josh Engels

3/6: Surprisingly, when we add this steering vector to different layers, the "in distribution" risky behavior and "out of distribution" self awareness are impacted identically! We think this means that the "awareness" mechanism is probably the same as the "behavior" mechanism.

Image in tweet by Josh Engels

4/6: We also study "risk backdoors": the LLM is trained to act risky only when a backdoor is present. Unfortunately, we don't reproduce the original paper's backdoor awareness results, but we do analyze the surprising fact that steering vectors can implement conditional logic!

Image in tweet by Josh Engels

5/6: Check out our post for more details! This is an interim progress report, so we're still looking into this; I'm very excited about more complex self awareness behaviors. Thanks to @NeelNanda5 and @sen_r for their always excellent collaboration. https://www.lesswrong.com/post...

6/6: At a higher level, I think that this is a good direction for mech interp: take some weird model behaviors and try to explain them. You can then step back and try to draw larger conclusions about what is going on in LLMs, and ideally develop new mech interp tools as a result.

Share this thread

Read on Twitter

View original thread

Navigate thread

1/6