VAIL
research area

Interpretability

Interpretability asks what a trained model has actually learned, rather than what we assume it learned from watching its accuracy go up. We work on attribution and saliency, concept probes, and mechanistic analysis of the representations inside vision and multimodal models, including how those representations transfer when a model is asked to do something it was not trained for.

We are equally interested in whether an explanation can be believed. A saliency map can look convincing and still fail to describe what the model did, so a fair amount of our work goes into testing explanations themselves instead of accepting them because they look reasonable.

Publications

  • FAILGROUND: Constructing and Safely Replaying LLM Agent Failure Knowledge Bases

    Anonymous · Under review · 2026

  • PRAXIS: Principled Reasoning via Agentic Exploration at Inference-time Scale

    Anonymous · Under review · 2026

  • When Agreement Is Not Enough: A Selection Bottleneck in Non-Verifiable Reasoning

    Shreshth Rai, Sarthak Pandey, Seifedine Kadry · Second Workshop on Curated Data for Efficient Learning @ ECCV · 2026