The year/Independent research

Paper 2605.23895

From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain

Published
May 2026
Research lab
Independent
Citations
0
GitHub
Not linked

01 In brief

Summary

BrainCause is an automated framework for causally discovering and validating visual concept representations in the human brain using fMRI.

It addresses the limitation of activation-based methods, which often identify false positives driven by correlated visual or semantic cues.

BrainCause constructs targeted stimulus sets with positive images, counterfactual edits (removing the target concept), and semantic negatives (correlated but distinct concepts), generated via language and vision models.

It uses an image-to-fMRI encoder to predict brain responses and ranks voxels by a causal score combining activation and specificity against negatives and edits.

Evaluated on the Natural Scenes Dataset across 260 concepts, BrainCause reduces false positive rates from 73.4% to 23% compared to activation-based discovery, while improving true positive rates.

It recovers known functional regions (faces, bodies, places, words) and discovers fine-grained representations (e.g., body parts, text types) with cross-subject consistency.

The framework also proposes follow-up experiments when measured data coverage is insufficient, and identifies limitations in current generative models.

02 From the paper

Abstract

Identifying which brain regions represent a visual concept in the human brain is a central challenge in neuroscience. Existing approaches have localized coarse functional regions (e.g., faces, places) through activation maximization, identifying regions that activate strongly for a target concept relative to other concepts. Yet strong activation alone does not establish that a region represents the concept itself, as responses may instead be driven by correlated visual or semantic cues. We introduce BrainCause, an automated framework that combines generative and brain models to synthesize controlled stimuli and validate neural representations through targeted causal testing. Given a query specifying a concept of interest, our framework constructs targeted stimulus sets comprising concept images, counterfactual edits that remove the target concept while preserving other image content, and images with candidate correlated distractors. It then uses an image-to-fMRI encoding model to predict brain responses and searches for representations that respond specifically to the target concept over correlated alternatives. BrainCause returns validated candidate representations and proposes follow-up fMRI experiments to further test or extend its discoveries. Our approach successfully recovers known functional localizations and identifies new candidate representations across dozens of concepts, validated on both predicted and measured fMRI data. Critically, we show that without causal validation, a large fraction of localizations would be false positives, confirming that activation alone is insufficient evidence of representation.