Vaishnavi Khindkar
I am a Master's student in Computer Vision (MSCV) at the School of Computer Science
(Robotics Institute), Carnegie Mellon University, where I work with Prof.
Andrej Risteski on the theory of generative modeling and sampling, and with
Prof. László
A. Jeni on object-centric 4D reconstruction. I am also a Machine Learning Engineering Intern at
Sync Labs, working on video generative AI
and lip-sync pipelines. Earlier at CMU, I worked with Prof.
Fernando De la Torre on diagnosing aerial object detectors with foundation generative models.
Before CMU, I spent several years at IIIT Hyderabad. As a Research Scientist at CVIT, I
worked with Prof. C. V.
Jawahar, Prof. Vineeth
N. Balasubramanian, and Prof.
Chetan Arora on autonomous-driving perception, pedestrian-intent prediction, and domain-adaptive
detection, work that led to publications at IROS 2024 and WACV 2022, one U.S. patent, and two
Indian patents. As an Applied AI Researcher at Product Labs, I also contributed to
Bhashini, the Government of India's national language-AI mission. I hold a B.E. in
Computer Science from Savitribai Phule Pune University.
Outside research, you'll usually find me at the piano, singing, swimming, running, meditating, or lost in
a good book, with a soft spot for
भक्ती संगीत,
अभंग,
and sci-fi (Iron Man, always). Giving back matters
to me too: I've mentored students through AI4ALL and taught schoolkids to code.
Research interests
Generative Modeling
Spatial Intelligence
3D & 4D Scene Understanding
World Models
Autonomous Driving
Robotics
Research
My path into research started with a question I couldn't let go of as an undergrad: what would the next
revolution in computing really look like? I was fascinated by the mathematics beneath how things
work, especially the structure and uncertainty behind intelligent behavior. Half-seriously and half
inspired by Iron Man's Jarvis, I became captivated by the idea of mind-reading computers: machines that could
grasp intent and understand the why behind what people do. That, to me, was the frontier
worth chasing.
That fascination with intent grew into a research program. At CVIT, I worked on reading
pedestrian intent and modeling object-scene interaction, using causal,
explainable models to predict not just what an agent will do, but why, alongside earlier
work on domain adaptation, spatio-temporal reasoning with graph networks, and generative augmentation,
all tied to the broader question of how machines can understand behavior in complex, changing environments.
Over time, this question pulled me one level deeper. To understand why agents act, we also need to
understand the worlds they act within: the geometry of a scene, how objects move, how interactions unfold,
and how these dynamics persist over time. This shift drew me toward computer vision, generative
modeling, and spatial intelligence, especially models that go beyond simply perceiving a scene,
and instead reconstruct how the world unfolds across space and time.
Right now I'm pursuing two complementary threads: with Prof.
Andrej Risteski, the theory of reward-tilted sampling, on how to steer generative
models toward desired outcomes in a principled way; and with
Prof. László
A. Jeni, Slot4D, an object-centric approach to 4D reconstruction that links low-level
geometric tracking to object-level scene understanding.
Across these projects, I am interested in a broader question: how can we build models that move beyond
perception alone: models that understand structure, dynamics, intent, and the physical worlds in
which intelligent behavior unfolds?