Skip to main content

In Brief

We show that apparent differences in what neural activity represents can reverse when language inputs are better matched, meaning that model design can strongly shape representational conclusions.

Abstract

Foundation-model features are increasingly used to ask what information neural activity represents, often by comparing prediction gains between nested encoding models. We show that such multimodal contrasts can change sign when only the conditioning predictor is reconstructed. Using fMRI from the Natural Scenes Dataset, DINOv2 visual features, and MPNet embeddings of MS COCO captions and Localized Narratives, a caption-narrative contrast in the additional predictive contribution of vision favors narratives when one short caption is compared with a long narrative (+0.012/+0.015 in Places), but favors captions after approximate word-count matching (-0.031/-0.023). The shift occurs across every measured ROI in both subjects and is driven primarily by differences in language-only prediction. Comparable contrasts also survive removal of image-specific content-word identity in several ROIs. These results show that nested neural contrasts do not identify represented content by themselves: predictor construction is part of the experimental design, and matched controls are required for representational claims.

Related Research

We compare representations in biological and artificial visual systems to understand what makes them similar, where they differ, and which biological constraints matter for learning and behavior.

Lucas Nadolskis · Galen Pogoncheff

All publications