Kruttidipta Samal
Introspective Machines
Abstract
As large language models evolve into autonomous agents, understanding what they “think” becomes an engineering necessity. This talk surveys the interpretability stack — from mechanistic internals (sparse autoencoders, residual stream probes) to runtime guardrails — and extends the same defense-in-depth principles to multi-agent communication protocols.
If you wish to modify any information or update your photo, please contact Web Chair Arief Wicaksana.
