Skip to main content

Active Vision & Sensorimotor Intelligence

Much of what we know about visual cortex comes from animals viewing controlled stimuli while keeping their heads still. Natural vision is different. During navigation, foraging, and pursuit, animals actively move their eyes, head, and body, continually changing both what they see and how the brain responds.

What We Study

We use recordings from freely moving mice and from real-world and virtual navigation experiments to study how visual input, movement, and behavioral state jointly shape cortical activity. We model these signals together and also study the sampling behavior itself, including gaze shifts and head–eye coordination.

We also compare biological and artificial visual systems. Mice and AI models can be tested on the same tasks, allowing us to ask where their behavior and internal representations agree or differ. Digital twins of visual cortex provide another way to test hypotheses that would be difficult to probe directly in the brain. We are also interested in event-driven and spiking approaches to efficient visual processing.

Current Directions

  • Behavior-dependent cortical dynamics. Neural activity during freely moving behavior and models that account for visual input, movement, and internal state.
  • Active sampling. Gaze shifts, head–eye coordination, and other strategies that determine what visual information is acquired.
  • Mouse versus AI. Comparing biological and artificial agents on the same visual tasks and representations.
  • Digital twins. Predictive models of visual cortex that can be used to test hypotheses about neural computation.
  • Event-driven and spiking vision. Efficient visual processing inspired by biological sensing and computation.

Team


Publications

We review computations that are engaged in ecological contexts, including active sensing, motion processing, scene analysis, distance estimation, and spatial perception.

We found that freely moving mice use multiple structured head-eye coordination motifs to shift gaze, including a Head-with-Eye motif that appears to reflect active visual orienting during natural behavior.

We introduce Mouse vs. AI, a public benchmark suite that unifies visual robustness, embodied foraging behavior, and neural alignment by evaluating artificial agents and mice in the same naturalistic 3D task.

We propose a novel temporal-digital architecture that encodes ANN weights as delays and activations as signal arrival times, enabling full ANN execution with temporal reuse, noise-tolerant summation, and hybrid memory, achieving up to 11× energy and 4× latency improvements over SNNs, and 3.5× energy savings over 8-bit digital systolic arrays.

We introduce a multimodal recurrent neural network that integrates gaze-contingent visual input with behavioral and temporal dynamics to explain V1 activity in freely moving mice.

We present a way to implement long short-term memory (LSTM) cells on spiking neuromorphic hardware.

We present a SNN model that uses spike-latency coding and winner-take-all inhibition to efficiently represent visual objects with as little as 15 spikes per neuron.

Featured Research

We developed a spiking neural network model that showed MSTd-like response properties can emerge from evolving spike-timing dependent plasticity with homeostatic synaptic scaling (STDP-H) parameters of the connections between area MT and MSTd.

We present a SNN model that uses spike-latency coding and winner-take-all inhibition to efficiently represent visual stimuli from the Fashion MNIST dataset.

Using a dimensionality reduction technique known as non-negative matrix factorization, we found that a variety of medial superior temporal (MSTd) neural response properties could be derived from MT-like input features. The responses that emerge from this technique, such as 3D translation and rotation selectivity, spiral tuning, and heading …

We present a cortical neural network model for visually guided navigation that has been embodied on a physical robot exploring a real-world environment. The model includes a rate based motion energy model for area V1, and a spiking neural network model for cortical area MT. The model generates a cortical representation of optic flow, determines the …

We present a two-stage model of visual area MT that we believe to be the first large-scale spiking network to demonstrate pattern direction selectivity. In this model, component-direction-selective (CDS) cells in MT linearly combine inputs from V1 cells that have spatiotemporal receptive fields according to the motion energy model of Simoncelli and …

Back to all research areas