VAIL
research area

AI Safety

A model that scores well in evaluation can still fail once it is deployed, and the failures that matter are rarely the ones a benchmark measures. We study how vision and language systems behave under distribution shift, how their failure modes can be found and written down before users run into them, and what happens when a system is given the chance to learn from its own past mistakes.

Some of this is robustness in the ordinary sense, where the question is how far a model can be pushed off its training distribution before it stops being reliable. Some of it concerns agents, where errors compound over long runs and one bad step can be repeated indefinitely if nothing catches it. The aim in both cases is to know the limits of a system before anyone has to depend on it.

Diagram of a pipeline that mines agent errors into a failure knowledge graph and replays them in later episodes.
Failure knowledge graph and replay pipeline from FAILGROUND, our work on agent failure memory.

Publications

  • FAILGROUND: Constructing and Safely Replaying LLM Agent Failure Knowledge Bases

    Anonymous · Under review · 2026

  • MedCurate-Bench: Auditing the Diagnostic Validity of Curated Medical Image Datasets

    Sarthak Pandey, Shreshth Rai, Seifedine Kadry · Second Workshop on Curated Data for Efficient Learning @ ECCV · 2026

  • PRAXIS: Principled Reasoning via Agentic Exploration at Inference-time Scale

    Anonymous · Under review · 2026

  • When Agreement Is Not Enough: A Selection Bottleneck in Non-Verifiable Reasoning

    Shreshth Rai, Sarthak Pandey, Seifedine Kadry · Second Workshop on Curated Data for Efficient Learning @ ECCV · 2026