Mar 28 2025 22 mins 2
Researchers used a novel "circuit tracing" method to explore how Claude 3.5 Haiku works internally. They mapped out how the model handles tasks like reasoning, poetry, translation, and math, identifying key features and how they interact. The study reveals complex strategies like planning and explores behaviors like hallucinations and refusals. Their findings offer new insights into how large models compute, aiming to make AI more interpretable and safer.
Podcast:
https://kabir.buzzsprout.com
YouTube:
https://www.youtube.com/@kabirtechdives
Please subscribe and share.