Interpretability
Interpretability asks what a trained model has actually learned, rather than what we assume it learned from watching its accuracy go up. We work on attribution and saliency, concept probes, and mechanistic analysis of the representations inside vision and multimodal models, including how those representations transfer when a model is asked to do something it was not trained for.
We are equally interested in whether an explanation can be believed. A saliency map can look convincing and still fail to describe what the model did, so a fair amount of our work goes into testing explanations themselves instead of accepting them because they look reasonable.
Publications
FAILGROUND: Constructing and Safely Replaying LLM Agent Failure Knowledge Bases
Anonymous · Under review · 2026
PRAXIS: Principled Reasoning via Agentic Exploration at Inference-time Scale
Anonymous · Under review · 2026
When Agreement Is Not Enough: A Selection Bottleneck in Non-Verifiable Reasoning
Shreshth Rai, Sarthak Pandey, Seifedine Kadry · Second Workshop on Curated Data for Efficient Learning @ ECCV · 2026