A foundation model for speech-evoked brain activity.

RABBiT predicts speech- and language-evoked fMRI responses directly from audio.

One model captures shared brain responses and adapts to individual subjects.

  • Zero-shot prediction

    Population-level responses to speech, without fMRI.

  • Few-shot adaptation

    Personalize to a novel subject with a few minutes of audio and fMRI.

  • Outperforms state of the art

    Better brain alignment in speech and language regions.

  • Language neuroscience

    Reproduce language findings in neuroscience.

Omer Moussa and Mariya Toneva

Two ways to use RABBiT

Choose the prediction setting based on the participant data available. To our knowledge, RABBiT is the only speech-to-fMRI foundation model that natively supports both zero-shot prediction and parameter-efficient few-shot adaptation.

Zero-shot

No fMRI needed

  1. AudioAny speech
  2. RABBiTNo fitting
  3. PredictionPopulation response

Predict population-level brain responses to any speech, without requiring fMRI.

What data is needed?

Only the speech audio is needed at prediction time. No participant-specific fitting is performed. The model is already pretrained on paired audio and fMRI.

Explore zero-shot results

Few-shot

A short calibration recording

  1. Audio + fMRIShort calibration
  2. Adapt RABBiTFit to participant
  3. PredictionParticipant-specific

Adapt predictions using a small amount of the participant’s fMRI and corresponding audio. Approximately 10 minutes of recordings improves participant-specific prediction in the evaluated setting.

What data is needed?

Paired audio and fMRI are used for calibration; subsequent predictions use new audio. Ten minutes refers to recording duration, not optimization time. Evaluation uses separate test recordings.

Explore few-shot results

Results

Prediction without participant data, adaptation with a short recording, and analyses of learned brain representations.

How to read the metrics

Group-level alignment

Pearson correlation between predicted and group-average fMRI over time, averaged over vertices within each region. This measures shared response structure.

Individual-level alignment

Correlation with a subject’s measured fMRI. Few-shot evaluation uses a separate test recording to assess adaptation to that subject.

Inter-subject consistency (ISC)

Correlation between one subject’s fMRI and the mean of the other subjects. This is a reference for shared responses, not a strict upper bound on group-level alignment.

Zero-shot results

Population prediction from speech alone.

Zero-shot evaluation uses held-out speech from Narratives and Le Petit Prince, covering 324 participants with no overlap with training. The comparisons below measure brain alignment with group-average fMRI responses.

Comparison with TRIBEv2

Cortical difference in correlation: RABBiT minus TRIBEv2. Blue favors RABBiT and orange favors TRIBEv2; the legend spans minus 0.2 to plus 0.2.
Fig 4Mean difference in Pearson correlation: Δr = RABBiT − TRIBEv2. Blue favors RABBiT; orange favors TRIBEv2. Display range: −0.2 to +0.2.

Largest gains in auditory and temporal cortex

RABBiT shows stronger brain alignment in auditory and temporal language regions. The models perform similarly across several higher-order language regions; TRIBEv2 leads in precuneus/PCC.

This map averages seven Narratives stories and eight Le Petit Prince sections, using group-average fMRI as the target.

How this comparison is measured
Prediction target
Pearson correlation over time between each model’s prediction and the mean fMRI response of subjects hearing the same recording, at each cortical vertex.
Aggregation
Compute RABBiT − TRIBEv2 for each recording, then take the arithmetic mean across the 15 included recordings. Le Petit Prince section 2 is absent from this figure’s export.
Listening inputs and surface
TRIBEv2’s visual input is zeroed. Its correlations on fsaverage5 are mapped to fsaverage6 by nearest-neighbor interpolation for this cortical visualization.
Temporal alignment
The exported map uses saved, model-specific timing offsets for each recording before correlation is computed. The offsets need not be identical across models or recordings.

Regional significance and cohort-level summaries follow below. See metric definitions for group-level alignment, individual responses, and inter-subject consistency.

Per-region zero-shot brain alignment: RABBiT vs TRIBEv2, with legend for the two models, linear baseline, ISC reference, and significance markers
Fig 5Per-region zero-shot brain alignment. RABBiT vs TRIBEv2.
Region by region

Performance varies by region.

Region by region, RABBiT's largest gains occur in auditory cortex, STG / STS, temporal pole, insula / FOp, and mPFC. The two models perform similarly across much of higher-order language cortex, with precuneus / PCC the one region where TRIBEv2 significantly leads.

Visual and motor regions remain near chance for both models in this naturalistic listening evaluation.

Group-level language-ROI correlation for RABBiT, TRIBEv2, linear and full-rank baselines, alongside peer and leave-one-out inter-subject consistency estimates
Fig 6$r_\text{group}$ across bilateral language ROIs. Narratives & Le Petit Prince.
vs the inter-subject consistency measure (ISC)

Group-level prediction and inter-subject consistency.

Across the two cohorts, RABBiT has higher language-ROI group-level correlation than TRIBEv2 and the linear baseline, with performance similar to the full-rank readout variant. Here, $r_\text{group}$ correlates the prediction with the cohort-mean fMRI response.

Gray dots show peer inter-subject consistency; the black dashed line is the leave-one-out ISC estimate, labeled a ceiling in the source figure. It provides a reference for shared response structure, rather than an absolute upper bound on model accuracy or evidence of complete prediction of an individual's response.

Few-shot results

Adaptation from a short calibration recording.

Few-shot adaptation uses a contiguous segment of paired audio and fMRI from the new participant, then evaluates predictions on a separate, fixed test segment. The main benchmark uses 19 participants listening to the 21st Year narrative and calibration recordings of 5–40 minutes. The ten-minute examples below refer to recording duration, not fitting time.

Few-shot adaptation: the fixed RABBiT encoder feeds shared and idiosyncratic coefficient heads, each paired with fixed cortical bases. Their responses sum to predicted fMRI. A few minutes of paired audio and fMRI from a new subject tune only the idiosyncratic head: about 115K fitted parameters, roughly 2,000 times fewer than voxel-wise ridge.
Few-shot predicted-vs-actual correlation for a held-out participant listening to the 21st Year narrative, after ~10 minutes of calibration recordings
Fig 7Held-out participant (21st Year narrative), using ~10 min of calibration recordings. Normalized 0–1 display scale. Fig 8 shows the change in correlation directly.
Few-shot accuracy

Prediction after participant-specific adaptation.

This example shows prediction after fitting the deviation coefficient maps using ~10 minutes of paired audio and fMRI. The backbone, temporal brain transformer, shared pathway, and deviation bases remain fixed.

Cortical map of few-shot minus zero-shot correlation for the same listener: red marks improvement across higher-order language regions, auditory cortex unchanged
Fig 8Δr = few-shot − zero-shot, same listener, −0.1 to +0.1. red ⇒ improved by the 10-min correction, blue ⇒ reduced.
Few-shot minus zero-shot

Regional changes after adaptation.

For this participant, the larger improvements occur in higher-order language regions, including frontal, inferior-parietal, medial frontal, and precuneus / PCC regions. Auditory cortex changes comparatively little.

Zero-shot and few-shot correlation in the displayed language regions after ten minutes of calibration recordings, with significance markers
Fig 9Per-region accuracy. zero-shot (blue) vs few-shot (red). ∗∗ p<0.01, ∗∗∗ p<0.001.
Regional comparison

Zero-shot and adapted predictions.

The displayed higher-order language regions show increased correlation after adaptation using ten minutes of recordings. Both bars use the same pretrained RABBiT model, before and after participant-specific fitting; markers indicate the reported significance levels.

Few-shot vs ridge regression as a function of calibration recording duration, with a legend for the four series: pretrained ridge, brain-tuned ridge, RABBiT without SID, and RABBiT
Fig 10Mean language-ROI correlation vs calibration recording duration. 19 participants, fixed test segment. markers are paired t-tests.
vs per-voxel ridge

Comparison with voxel-wise ridge regression.

On the main benchmark, RABBiT improves over its zero-shot prediction with 5 minutes of calibration recordings and exceeds both ridge baselines at the tested durations from 5 to 40 minutes. Each method uses the same participant's calibration data.

RABBiT updates ~115K parameters per participant, compared with approximately 222M for the wav2vec2 voxel-wise ridge readout described in the appendix. These are fitted adaptation/readout parameters, not total model sizes.

Per-region relative change in correlation from zero-shot at ten minutes of calibration recordings, showing pretrained ridge, brain-tuned ridge, RABBiT without SID, and RABBiT
Fig 11Improvement over own zero-shot at ~10 min. ∗∗ / ∗∗∗ vs zero-shot.
Ten-minute calibration recordings

Relative improvement over zero-shot.

At approximately ten minutes of calibration data, the paper reports 30–90% relative increases in correlation in the displayed higher-order regions. The reference is RABBiT's own zero-shot prediction, not ridge regression; these are relative changes, not percentage-point gains.

Reproducing neuroscience findings

Probing auditory and language representations.

In addition to prediction accuracy, we examine the model's responses to a speech-intelligibility contrast and the organization of its learned region representations. These analyses test correspondence with established properties of auditory and language cortex.

Predicted intact-minus-degraded speech contrast, with a left-lateralized pattern in language-associated cortex
Fig 1Predicted intact − degraded speech on the cortical surface.
Fedorenko language localizer

Intelligible vs unintelligible speech localizer

Subtracting predicted responses to acoustically degraded speech from those to intact speech yields a left-lateralized pattern in frontal, lateral temporal, inferior-parietal, and angular regions.

The predicted contrast is smaller in early auditory cortex. This analysis probes a model trained on naturalistic stimuli, without fitting it to the localizer contrast.

A greedy similarity walk through the model's learned per-region representations traces the auditory-to-language processing hierarchy
Fig 2Greedy similarity walk through the model's learned ROI representations.
Emergent hierarchy

A coarse auditory-to-language progression.

Starting from primary auditory cortex, a greedy walk to the nearest unvisited ROI query in cosine-similarity space passes through belt regions, STG / STS, temporal association cortex, and parietal and frontal language regions.

No hierarchy labels supervise these queries. The resulting ordering is consistent with a coarse auditory-to-language progression; it does not establish a causal processing pathway.

These analyses characterize the model's predictions and learned representations. The zero-shot results evaluate prediction accuracy without participant-specific fitting; the few-shot results evaluate adaptation using paired audio and fMRI recordings.

Model architecture

How RABBiT maps speech to shared and participant-specific brain responses.

The Temporal Brain Transformer constructs region-level representations from speech. Shared–Idiosyncratic Decomposition maps these representations to a shared response component and a participant-specific deviation. Together they support population prediction and adaptation of the deviation pathway.

TBT Temporal Brain Transformer

Region-specific attention over the speech stream.

RABBiT learns one query per cortical region. Each query attends across the speech sequence to produce a region-level representation $z_i$. The model uses 30 regions (15 bilateral HCP-MMP1 groupings on fsaverage6); the SID readout then maps these representations to vertex-level predictions. The expression below summarizes a single attention head; the full model uses stacked multi-head blocks.

Attention equation
$$\alpha_{i,t} = \operatorname{softmax}_t \! \left( \langle W_Q\, q_i, \; W_K\, x_t \rangle \right), \qquad z_i = \sum_t \alpha_{i,t} \cdot W_V\, x_t \;\in\; \mathbb{R}^{d_o}$$
RABBiT cross-attention readout: ROI query tokens attend to the speech-token sequence
Fig 12ROI query tokens attend to the speech-token sequence.
TBT readout

One query per region.

Each brain region learns its own query over the speech stream. Rather than pooling speech uniformly, every region selectively gathers the moments most relevant to it, producing a compact region-level representation zi.

SID Shared–Idiosyncratic Decomposition

Shared population structure and participant-specific deviations.

Each region's predicted response combines a shared basis ($\Phi_i$) and a participant-specific deviation basis ($\Delta_{i,s}$), with stimulus-dependent coefficients. For a new participant, zero-shot prediction uses the shared basis together with the average deviation basis learned from the training cohort. Few-shot adaptation retains those bases and fits the deviation coefficient maps to the new participant's calibration recordings.

Prediction equation
$$\hat{y}^{(s)}_i = \pi_i(z_i) \, \Phi_i \;+\; \rho_i(z_i) \, \Delta_{i,s}$$

$\pi_i$ and $\rho_i$ map the region representation $z_i$ to shared and deviation coefficients. For a new participant, $\Delta_{i,s}$ is initialized to the training-cohort average and held fixed. Few-shot updates only the parameters of $\rho_i$; the backbone, transformer, shared maps, and all bases remain frozen.

RABBiT pipeline: audio → speech backbone → TBT → SID → predicted fMRI
Fig 13audio → speech backbone → TBT → SID → predicted fMRI.
End to end

From speech to brain.

A brain-tuned speech model converts audio into speech representations. The Temporal Brain Transformer routes them into region-level representations. SID then separates every prediction into a shared population response and a compact participant-specific deviation before producing high-resolution fMRI predictions.

Training uses paired audio and fMRI from CNeuroMod Friends, with LoRA adaptation of the speech backbone. Cohort size, recording conditions, and data splits are documented in the paper's Methods and appendix.

Component analyses and ablations

Evaluating model components and shared response structure.

Component ablations: removing brain-tuning or the temporal brain transformer hurts; the decomposition matches a far larger per-subject readout
Fig 14Zero-shot accuracy as components are removed or swapped.
Ablations

Contributions to zero-shot prediction.

Removing brain-tuning produces the largest reduction in zero-shot correlation; removing the Temporal Brain Transformer also reduces performance. Replacing SID with full participant-specific readouts gives no significant improvement in the reported comparison, despite roughly five times as many trainable model parameters. Replacing wav2vec2-base with WavLM-large provides little additional benefit in this evaluation.

Correlation with a group average including or excluding the participant, compared across auditory and language regions
Fig 15Group average including (blue) vs excluding (orange) the listener.
The diagnostic

Shared and participant-specific response structure.

Correlations with the group average decrease more in higher-order language regions when the participant is excluded from that average. This pattern motivates the distinction between shared and participant-specific response components.

The same regions tend to show larger few-shot gains. The averaging diagnostic describes this association; it does not by itself establish the cause of the gains.

Evaluation scope. The reported transfer results primarily concern English naturalistic listening and auditory/language regions. Generalization to other languages, reading, conversation, and multimodal tasks requires further evaluation. Few-shot results concern adaptation to individual participants, rather than a single calibration for an entire new dataset. Training and evaluation details are provided in the paper's Methods and limitations.

Resources

Code, weights, demos.

Paper, implementation, prediction examples, and source datasets.

Model availability

The RABBiT model repository currently provides the rabbit_fp32.onnx browser export. Its approximately 422 MB download starts only when you choose to load the microphone model. Recorded examples work without these weights.

Source datasets

Citation

BibTeX

If RABBiT is useful for your work, please cite it.

Download .bib
@article{moussa2026rabbit,
  title     = {RABBiT: Rapidly Adaptive {BOLD} Foundation Model via Brain-Tuning
               for Accurate Zero-Shot and Few-Shot Prediction of
               Speech-Elicited Responses in the Brain},
  author    = {Moussa, Omer and Toneva, Mariya},
  journal   = {arXiv preprint arXiv:2607.05171},
  year      = {2026},
  doi       = {10.48550/arXiv.2607.05171},
  url       = {https://arxiv.org/abs/2607.05171}
}

Acknowledgements

Acknowledgements. RABBiT builds on data from the CNeuroMod project, the Narratives dataset, and Le Petit Prince (Bhattasali et al.). We thank the CNeuroMod team for the Friends fMRI release and the TRIBEv2 authors for code + checkpoints.