Choose the prediction setting based on the participant data available. To our knowledge, RABBiT is the only speech-to-fMRI foundation model that natively supports both zero-shot prediction and parameter-efficient few-shot adaptation.
Zero-shot
No fMRI needed
AudioAny speech
RABBiTNo fitting
PredictionPopulation response
Predict population-level brain responses to any speech, without requiring fMRI.
What data is needed?
Only the speech audio is needed at prediction time. No participant-specific fitting is performed. The model is already pretrained on paired audio and fMRI.
Adapt predictions using a small amount of the participant’s fMRI and corresponding audio. Approximately 10 minutes of recordings improves participant-specific prediction in the evaluated setting.
What data is needed?
Paired audio and fMRI are used for calibration; subsequent predictions use new audio. Ten minutes refers to recording duration, not optimization time. Evaluation uses separate test recordings.
Prediction without participant data, adaptation with a short recording, and analyses of learned brain representations.
How to read the metrics
Group-level alignment
Pearson correlation between predicted and group-average fMRI over time, averaged over vertices within each region. This measures shared response structure.
Individual-level alignment
Correlation with a subject’s measured fMRI. Few-shot evaluation uses a separate test recording to assess adaptation to that subject.
Inter-subject consistency (ISC)
Correlation between one subject’s fMRI and the mean of the other subjects. This is a reference for shared responses, not a strict upper bound on group-level alignment.
Zero-shot results
Population prediction from speech alone.
Zero-shot evaluation uses held-out speech from Narratives and Le Petit Prince, covering 324 participants with no overlap with training. The comparisons below measure brain alignment with group-average fMRI responses.
Comparison with TRIBEv2
Fig 4Mean difference in Pearson correlation: Δr = RABBiT − TRIBEv2. Blue favors RABBiT; orange favors TRIBEv2. Display range: −0.2 to +0.2.
Largest gains in auditory and temporal cortex
RABBiT shows stronger brain alignment in auditory and temporal language regions. The models perform similarly across several higher-order language regions; TRIBEv2 leads in precuneus/PCC.
This map averages seven Narratives stories and eight Le Petit Prince sections, using group-average fMRI as the target.
How this comparison is measured
Prediction target
Pearson correlation over time between each model’s prediction and the mean fMRI response of subjects hearing the same recording, at each cortical vertex.
Aggregation
Compute RABBiT − TRIBEv2 for each recording, then take the arithmetic mean across the 15 included recordings. Le Petit Prince section 2 is absent from this figure’s export.
Listening inputs and surface
TRIBEv2’s visual input is zeroed. Its correlations on fsaverage5 are mapped to fsaverage6 by nearest-neighbor interpolation for this cortical visualization.
Temporal alignment
The exported map uses saved, model-specific timing offsets for each recording before correlation is computed. The offsets need not be identical across models or recordings.
Regional significance and cohort-level summaries follow below. See metric definitions for group-level alignment, individual responses, and inter-subject consistency.
Fig 5Per-region zero-shot brain alignment. RABBiT vs TRIBEv2.
Region by region
Performance varies by region.
Region by region, RABBiT's largest gains occur in auditory cortex, STG / STS, temporal pole, insula / FOp, and mPFC. The two models perform similarly across much of higher-order language cortex, with precuneus / PCC the one region where TRIBEv2 significantly leads.
Visual and motor regions remain near chance for both models in this naturalistic listening evaluation.
Fig 6$r_\text{group}$ across bilateral language ROIs. Narratives & Le Petit Prince.
vs the inter-subject consistency measure (ISC)
Group-level prediction and inter-subject consistency.
Across the two cohorts, RABBiT has higher language-ROI group-level correlation than TRIBEv2 and the linear baseline, with performance similar to the full-rank readout variant. Here, $r_\text{group}$ correlates the prediction with the cohort-mean fMRI response.
Gray dots show peer inter-subject consistency; the black dashed line is the leave-one-out ISC estimate, labeled a ceiling in the source figure. It provides a reference for shared response structure, rather than an absolute upper bound on model accuracy or evidence of complete prediction of an individual's response.
Few-shot results
Adaptation from a short calibration recording.
Few-shot adaptation uses a contiguous segment of paired audio and fMRI from the new participant, then evaluates predictions on a separate, fixed test segment. The main benchmark uses 19 participants listening to the 21st Year narrative and calibration recordings of 5–40 minutes. The ten-minute examples below refer to recording duration, not fitting time.
Fig 7Held-out participant (21st Year narrative), using ~10 min of calibration recordings. Normalized 0–1 display scale. Fig 8 shows the change in correlation directly.
Few-shot accuracy
Prediction after participant-specific adaptation.
This example shows prediction after fitting the deviation coefficient maps using ~10 minutes of paired audio and fMRI. The backbone, temporal brain transformer, shared pathway, and deviation bases remain fixed.
Fig 8Δr = few-shot − zero-shot, same listener, −0.1 to +0.1. red ⇒ improved by the 10-min correction, blue ⇒ reduced.
Few-shot minus zero-shot
Regional changes after adaptation.
For this participant, the larger improvements occur in higher-order language regions, including frontal, inferior-parietal, medial frontal, and precuneus / PCC regions. Auditory cortex changes comparatively little.
The displayed higher-order language regions show increased correlation after adaptation using ten minutes of recordings. Both bars use the same pretrained RABBiT model, before and after participant-specific fitting; markers indicate the reported significance levels.
Fig 10Mean language-ROI correlation vs calibration recording duration. 19 participants, fixed test segment. markers are paired t-tests.
vs per-voxel ridge
Comparison with voxel-wise ridge regression.
On the main benchmark, RABBiT improves over its zero-shot prediction with 5 minutes of calibration recordings and exceeds both ridge baselines at the tested durations from 5 to 40 minutes. Each method uses the same participant's calibration data.
RABBiT updates ~115K parameters per participant, compared with approximately 222M for the wav2vec2 voxel-wise ridge readout described in the appendix. These are fitted adaptation/readout parameters, not total model sizes.
Fig 11Improvement over own zero-shot at ~10 min. ∗∗ / ∗∗∗ vs zero-shot.
Ten-minute calibration recordings
Relative improvement over zero-shot.
At approximately ten minutes of calibration data, the paper reports 30–90% relative increases in correlation in the displayed higher-order regions. The reference is RABBiT's own zero-shot prediction, not ridge regression; these are relative changes, not percentage-point gains.
Reproducing neuroscience findings
Probing auditory and language representations.
In addition to prediction accuracy, we examine the model's responses to a speech-intelligibility contrast and the organization of its learned region representations. These analyses test correspondence with established properties of auditory and language cortex.
Fig 1Predicted intact − degraded speech on the cortical surface.
Fedorenko language localizer
Intelligible vs unintelligible speech localizer
Subtracting predicted responses to acoustically degraded speech from those to intact speech yields a left-lateralized pattern in frontal, lateral temporal, inferior-parietal, and angular regions.
The predicted contrast is smaller in early auditory cortex. This analysis probes a model trained on naturalistic stimuli, without fitting it to the localizer contrast.
Fig 2Greedy similarity walk through the model's learned ROI representations.
Emergent hierarchy
A coarse auditory-to-language progression.
Starting from primary auditory cortex, a greedy walk to the nearest unvisited ROI query in cosine-similarity space passes through belt regions, STG / STS, temporal association cortex, and parietal and frontal language regions.
No hierarchy labels supervise these queries. The resulting ordering is consistent with a coarse auditory-to-language progression; it does not establish a causal processing pathway.
These analyses characterize the model's predictions and learned representations. The zero-shot results evaluate prediction accuracy without participant-specific fitting; the few-shot results evaluate adaptation using paired audio and fMRI recordings.
Model architecture
How RABBiT maps speech to shared and participant-specific brain responses.
The Temporal Brain Transformer constructs region-level representations from speech. Shared–Idiosyncratic Decomposition maps these representations to a shared response component and a participant-specific deviation. Together they support population prediction and adaptation of the deviation pathway.
TBT
Temporal Brain Transformer
Region-specific attention over the speech stream.
RABBiT learns one query per cortical region. Each query attends across the speech sequence to produce a region-level representation $z_i$. The model uses 30 regions (15 bilateral HCP-MMP1 groupings on fsaverage6); the SID readout then maps these representations to vertex-level predictions. The expression below summarizes a single attention head; the full model uses stacked multi-head blocks.
Fig 12ROI query tokens attend to the speech-token sequence.
TBT readout
One query per region.
Each brain region learns its own query over the speech stream. Rather than pooling speech uniformly, every region selectively gathers the moments most relevant to it, producing a compact region-level representation zi.
SID
Shared–Idiosyncratic Decomposition
Shared population structure and participant-specific deviations.
Each region's predicted response combines a shared basis ($\Phi_i$) and a participant-specific deviation basis ($\Delta_{i,s}$), with stimulus-dependent coefficients. For a new participant, zero-shot prediction uses the shared basis together with the average deviation basis learned from the training cohort. Few-shot adaptation retains those bases and fits the deviation coefficient maps to the new participant's calibration recordings.
$\pi_i$ and $\rho_i$ map the region representation $z_i$ to shared and deviation coefficients. For a new participant, $\Delta_{i,s}$ is initialized to the training-cohort average and held fixed. Few-shot updates only the parameters of $\rho_i$; the backbone, transformer, shared maps, and all bases remain frozen.
A brain-tuned speech model converts audio into speech representations. The Temporal Brain Transformer routes them into region-level representations. SID then separates every prediction into a shared population response and a compact participant-specific deviation before producing high-resolution fMRI predictions.
Training uses paired audio and fMRI from CNeuroMod Friends, with LoRA adaptation of the speech backbone. Cohort size, recording conditions, and data splits are documented in the paper's Methods and appendix.
Component analyses and ablations
Evaluating model components and shared response structure.
Fig 14Zero-shot accuracy as components are removed or swapped.
Ablations
Contributions to zero-shot prediction.
Removing brain-tuning produces the largest reduction in zero-shot correlation; removing the Temporal Brain Transformer also reduces performance. Replacing SID with full participant-specific readouts gives no significant improvement in the reported comparison, despite roughly five times as many trainable model parameters. Replacing wav2vec2-base with WavLM-large provides little additional benefit in this evaluation.
Fig 15Group average including (blue) vs excluding (orange) the listener.
The diagnostic
Shared and participant-specific response structure.
Correlations with the group average decrease more in higher-order language regions when the participant is excluded from that average. This pattern motivates the distinction between shared and participant-specific response components.
The same regions tend to show larger few-shot gains. The averaging diagnostic describes this association; it does not by itself establish the cause of the gains.
Evaluation scope. The reported transfer results primarily concern English naturalistic listening and auditory/language regions. Generalization to other languages, reading, conversation, and multimodal tasks requires further evaluation. Few-shot results concern adaptation to individual participants, rather than a single calibration for an entire new dataset. Training and evaluation details are provided in the paper's Methods and limitations.
Resources
Code, weights, demos.
Paper, implementation, prediction examples, and source datasets.
The RABBiT model repository currently provides the rabbit_fp32.onnx browser export. Its approximately 422 MB download starts only when you choose to load the microphone model. Recorded examples work without these weights.
@article{moussa2026rabbit,
title = {RABBiT: Rapidly Adaptive {BOLD} Foundation Model via Brain-Tuning
for Accurate Zero-Shot and Few-Shot Prediction of
Speech-Elicited Responses in the Brain},
author = {Moussa, Omer and Toneva, Mariya},
journal = {arXiv preprint arXiv:2607.05171},
year = {2026},
doi = {10.48550/arXiv.2607.05171},
url = {https://arxiv.org/abs/2607.05171}
}