Research Engineer, Interpretability at Anthropic

San Francisco, CA (On-Site) · About the role: When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understan…