{"id":66396,"date":"2025-10-10T18:52:11","date_gmt":"2025-10-10T18:52:11","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/66396\/"},"modified":"2025-10-10T18:52:11","modified_gmt":"2025-10-10T18:52:11","slug":"temporal-recurrence-as-a-general-mechanism-to-explain-neural-responses-in-the-auditory-system","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/66396\/","title":{"rendered":"Temporal recurrence as a general mechanism to explain neural responses in the auditory system"},"content":{"rendered":"<p>We begin this section with a summary of common models of auditory neural responses and introduce a novel, Transformer-based architecture, as well as our fully recurrent model called StateNet. Then, we describe the electrophysiology datasets used in this study and set the mathematical framework of the neural response fitting task. Because conventional models and most datasets were already described in a prior study from our group<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Ran&#xE7;on, U., Masquelier, T. &amp; Cottereau, B. R. A general model unifying the adaptive, transient and sustained properties of ON and OFF auditory neural responses. PLoS Comput. Biol. 20, e1012288 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR67\" id=\"ref-link-section-d61703220e2280\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>, we invite readers to consult it for a more precise description of conventional models and datasets. Lastly, we explain our process for reverse-engineering models, which generalizes STRFs for arbitrary network depth, width, and degree of nonlinearity.<\/p>\n<p>Canonical computational models of auditory neural responses<\/p>\n<p>The aim of computational models of neural responses in auditory cortex is to convert (&#8220;encode&#8221;) incoming sound stimuli into time-varying firing rates\/probabilities that predict electrophysiological measurements made in auditory areas. Traditionally, these models use the cochleagram of the stimulus\u2014a spectrogram-like representation that mimics processing in the cochlea\u2014and are rate-based (as opposed to spiking). Many such models have been proposed in the literature, ranging from simple Linear (&#8220;STRF&#8221; or \u201cL&#8221;) approaches<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Aertsen, A. &amp; Johannesma, P. The spectro-temporal receptive field-a functional characteristic of auditory neurons. Biol. Cybern. 42, 133&#x2013;43 (1981).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR17\" id=\"ref-link-section-d61703220e2291\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Theunissen, F. E., Sen, K. &amp; Doupe, A. J. Spectral-temporal receptive fields of nonlinear auditory neurons obtained using natural sounds. J. Neurosci. 20, 2315&#x2013;2331 (2000).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR18\" id=\"ref-link-section-d61703220e2294\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a> to more complex methods based on multi-layer CNNs<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Pennington, J. R. &amp; David, S. V. A convolutional neural network provides a generalizable model of natural sound coding by neural populations in auditory cortex. PLoS Comput. Biol. 19, e1011110 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR31\" id=\"ref-link-section-d61703220e2298\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>. However, all these models use temporal convolutions with finite window lengths and, therefore, finite TRFs. In this case, the duration of the TRF is a hyperparameter that is arbitrarily defined by the modeling scientist.<\/p>\n<p>More formally, a model \\({{{\\mathcal{M}}}}\\) is a causal application \\({{\\mathbb{R}}}^{F\\times T}\\mapsto {{\\mathbb{R}}}^{N\\times T}\\) where F is the number of frequency bands of a stimulus spectrogram, T a variable number of time steps, and N a number of units\/channels whose activity to predict.<\/p>\n<p>In this paper, we use five models based on this approach: the Linear (L) model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Aertsen, A. &amp; Johannesma, P. The spectro-temporal receptive field-a functional characteristic of auditory neurons. Biol. Cybern. 42, 133&#x2013;43 (1981).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR17\" id=\"ref-link-section-d61703220e2389\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Theunissen, F. E., Sen, K. &amp; Doupe, A. J. Spectral-temporal receptive fields of nonlinear auditory neurons obtained using natural sounds. J. Neurosci. 20, 2315&#x2013;2331 (2000).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR18\" id=\"ref-link-section-d61703220e2392\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>, the Linear-Nonlinear (LN) model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 45\" title=\"Chichilnisky, E. A simple white noise analysis of neuronal light responses. Network 12, 199&#x2013;213 (2001).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR45\" id=\"ref-link-section-d61703220e2396\" rel=\"nofollow noopener\" target=\"_blank\">45<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 68\" title=\"Ahrens, M. B., Linden, J. F. &amp; Sahani, M. Nonlinearities and contextual influences in auditory cortical responses modeled with multilinear spectrotemporal methods. J. Neurosci. 28, 1929&#x2013;1942 (2008).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR68\" id=\"ref-link-section-d61703220e2399\" rel=\"nofollow noopener\" target=\"_blank\">68<\/a>, the Network Receptive Field (NRF) model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Harper, N. S. et al. Network receptive field modeling reveals extensive integration and multi-feature selectivity in auditory cortical neurons. PLoS Comput. Biol. 12, e1005113 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR22\" id=\"ref-link-section-d61703220e2403\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>, the Dynamic Network (DNet) model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Rahman, M., Willmore, B. D. B., King, A. J. &amp; Harper, N. S. A dynamic network model of temporal receptive fields in primary auditory cortex. PLoS Comput. Biol. 15, e1006618 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR25\" id=\"ref-link-section-d61703220e2407\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>, and finally a deep 2D-CNN model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Pennington, J. R. &amp; David, S. V. A convolutional neural network provides a generalizable model of natural sound coding by neural populations in auditory cortex. PLoS Comput. Biol. 19, e1011110 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR31\" id=\"ref-link-section-d61703220e2411\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>. A general architecture of a convolutional model is illustrated in Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>a. Because a full description of these approaches was already provided in previous benchmarks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Ran&#xE7;on, U., Masquelier, T. &amp; Cottereau, B. R. A general model unifying the adaptive, transient and sustained properties of on and off auditory neural responses. PLoS Comput. Biol. 20, 1&#x2013;32 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR27\" id=\"ref-link-section-d61703220e2419\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>, we only review here some of their shared general features. We also report minor modifications that we introduced in their implementations so as to make them work on our unified PyTorch pipeline. We invite the reader to consult the original studies that introduced these models for a more detailed description of their functioning.<\/p>\n<p>Auditory periphery<\/p>\n<p>The initial processing stage converts the stimulus sound waveform into a biologically plausible spectrogram-like representation \\(x\\in {{\\mathbb{R}}}^{F\\times T}\\), thereby reflecting operations realized by the cochlea. In the literature, the waveform-to-spectrogram transformation can be performed through a simple short-term Fourier decomposition, or more often through temporal convolutions with a bank of mel or gammatone filters that are scaled logarithmically along the frequency axis. Following the latter, a compressive function such as a cubic root or logarithm is applied. Although the combination of both of these operations makes a consensus, there is a variability across studies in their implementation. However, it was shown that such variations are all more or less equivalent and still provide good cochlear sound encodings when modeling higher-order auditory neural responses. As a result, simple transformations should be preferred<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 69\" title=\"Rahman, M., Willmore, B. D. B., King, A. J. &amp; Harper, N. S. Simple transformations capture auditory input to cortex. Proc. Natl. Acad. Sci. USA 117, 28442&#x2013;28451 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR69\" id=\"ref-link-section-d61703220e2468\" rel=\"nofollow noopener\" target=\"_blank\">69<\/a>. In order to facilitate present and future comparisons with previous methods, and to limit as much as possible the introduction of biases due to different data pre-processings, we directly use here the cochleagrams provided in each dataset.<\/p>\n<p>Core principle<\/p>\n<p>Classical models rely on a cascade of temporal convolutions with a stride of 1 performed on the cochleagram of the sound stimulus, interleaved with standard nonlinear activation functions (e.g., Sigmoid, LeakyReLU) and followed by a parametric output nonlinearity with learnable parameters (e.g., baseline activity, slope saturation value). In all models, the cochleagram is systematically padded to the left (i.e., in the past) with zeroes prior to the temporal convolution operations, in order to respect causality and to output a time series of neural activity with as many time bins as in the input cochleagram.<\/p>\n<p>Single unit vs. population fitting<\/p>\n<p>In datasets where all sensory neurons were probed using the same set of stimuli, it is possible for computational models to predict the (vector) activity of the whole population<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Pennington, J. R. &amp; David, S. V. A convolutional neural network provides a generalizable model of natural sound coding by neural populations in auditory cortex. PLoS Comput. Biol. 19, e1011110 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR31\" id=\"ref-link-section-d61703220e2489\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>. This population coding paradigm allows to train a single model with some learnable parameters shared across all neural units under consideration, and some specific to each unit. As a result, the common backbone tends to learn robust and meaningful embeddings, which further reduces overfitting. Performances are, on average, better across the population than when fitting an entire model for each unit. Furthermore, this process drastically reduces training time and brings it down to a computational complexity of O(1) instead of O(N), where N is the total number of units in each dataset. For these reasons, we adopt the population coding paradigm whenever possible, that is, when various single-unit responses were recorded for the same stimuli. This is the case for the NS1, NAT4-A1, NAT4-PEG, AA1-MLd and AA1-Field_L datasets.<\/p>\n<p>Output nonlinearity<\/p>\n<p>All but the L model were equipped with a parametric nonlinear output activation function, learned alongside all other parameters through gradient descent. We used the following 4-parameter double exponential:<\/p>\n<p>$$f(x)=a\\exp (-\\exp (kx-s))+b$$<\/p>\n<p>\n                    (1)\n                <\/p>\n<p>where b represents the baseline spike rate, a the saturated firing rate, s the firing threshold, and k the gain<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Thorson, I. L., Li&#xE9;nard, J. &amp; David, S. V. The essential complexity of auditory receptive fields. PLoS Comput. Biol. 11, e1004628 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR21\" id=\"ref-link-section-d61703220e2602\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Pennington, J. R. &amp; David, S. V. A convolutional neural network provides a generalizable model of natural sound coding by neural populations in auditory cortex. PLoS Comput. Biol. 19, e1011110 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR31\" id=\"ref-link-section-d61703220e2605\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>. Importantly, in the case of population models (i.e., when predicting the simultaneous activity of several units, see below), each output neuron learns a different set of these four parameters.<\/p>\n<p>Regularization and parameterization<\/p>\n<p>Canonical models based on convolutions are prone to overfitting, and many strategies were proposed to limit this effect, such as the parameterization of spectro-temporal convolutional kernels<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Thorson, I. L., Li&#xE9;nard, J. &amp; David, S. V. The essential complexity of auditory receptive fields. PLoS Comput. Biol. 11, e1004628 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR21\" id=\"ref-link-section-d61703220e2617\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 60\" title=\"Woolley, S. M. N., Gill, P. R. &amp; Theunissen, F. E. Stimulus-dependent auditory tuning results in synchronous population coding of vocalizations in the songbird midbrain. J. Neurosci. 26, 2499&#x2013;2512 (2006).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR60\" id=\"ref-link-section-d61703220e2620\" rel=\"nofollow noopener\" target=\"_blank\">60<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 68\" title=\"Ahrens, M. B., Linden, J. F. &amp; Sahani, M. Nonlinearities and contextual influences in auditory cortical responses modeled with multilinear spectrotemporal methods. J. Neurosci. 28, 1929&#x2013;1942 (2008).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR68\" id=\"ref-link-section-d61703220e2623\" rel=\"nofollow noopener\" target=\"_blank\">68<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 70\" title=\"Willmore, B. &amp; Smyth, D. Methods for first-order kernel estimation: simple-cell receptive fields from responses to natural scenes. Network 14, 553&#x2013;77 (2003).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR70\" id=\"ref-link-section-d61703220e2626\" rel=\"nofollow noopener\" target=\"_blank\">70<\/a>. To stick to the most extensively reviewed version of these canonical models, as well as to highlight their limitations, our implementation did not include any such methods. Furthermore, we did not use any data augmentation techniques, weight decay, or dropout during training, as it was previously shown that such approaches complexify training and yield little to no improvements in performances<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Ran&#xE7;on, U., Masquelier, T. &amp; Cottereau, B. R. A general model unifying the adaptive, transient and sustained properties of on and off auditory neural responses. PLoS Comput. Biol. 20, 1&#x2013;32 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR27\" id=\"ref-link-section-d61703220e2630\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Pennington, J. R. &amp; David, S. V. A convolutional neural network provides a generalizable model of natural sound coding by neural populations in auditory cortex. PLoS Comput. Biol. 19, e1011110 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR31\" id=\"ref-link-section-d61703220e2633\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>. Instead, we used Batch Normalization (BN), which greatly improved the robustness and performances of all models, including the linear one (L), without compromising its nature, as after training, BN\u2019s scale and bias terms can be absorbed by the models own weights and biases.<\/p>\n<p>Transformer model<\/p>\n<p>Over the last years, attention-based Transformer architectures<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Vaswani, A. et al. Attention is all you need. In Proc. Advance Neural Information Processing Systems (eds Guyon, I. et al.) Vol. 30 &#010;                  https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf&#010;                  &#010;                . (Curran Associates, Inc., 2017).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR47\" id=\"ref-link-section-d61703220e2646\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a> have been more and more used by the AI community as an alternative to RNN for modeling long sequences<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 57\" title=\"Zhou, H. et al. Informer: beyond efficient transformer for long sequence time-series forecasting. In Proc. The Thirty-Fifth AAAI Conference on Artificial Intelligence, Virtual Conference Vol. 35, 11106&#x2013;11115 (AAAI Press, 2021).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR57\" id=\"ref-link-section-d61703220e2650\" rel=\"nofollow noopener\" target=\"_blank\">57<\/a>, from text<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 56\" title=\"Devlin, J., Chang, M.-W., Lee, K. &amp; Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C. &amp; Solorio, T. (eds.). In Proc. Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Vol 1 4171&#x2013;4186 &#010;                  https:\/\/aclanthology.org\/N19-1423&#010;                  &#010;                (Long and Short Papers).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR56\" id=\"ref-link-section-d61703220e2654\" rel=\"nofollow noopener\" target=\"_blank\">56<\/a> to images<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 55\" title=\"Dosovitskiy, A. et al. An image is worth 16x16 words: transformers for image recognition at scale. In Proc. International Conference on Learning Representations &#010;                  https:\/\/openreview.net\/forum?id=YicbFdNTTy&#010;                  &#010;                (2021).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR55\" id=\"ref-link-section-d61703220e2658\" rel=\"nofollow noopener\" target=\"_blank\">55<\/a>. Contrary to stateful approaches, which rely on BPTT for training, Transformers do not suffer from vanishing or exploding gradients. However, they present some drawbacks, such as quadratic algorithmic complexity scaling with the sequence length. To investigate whether this model is well-suited for fitting dynamic neural responses in the auditory cortex, we developed a novel architecture based on the attention mechanism (see Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>b). To the best of our knowledge, it is the first model of its kind that is proposed for this task.<\/p>\n<p>As for stateless models, a hyperparameter T defines the length of the temporal context window that serves to predict single-unit or population activity at the current time step. Within this window, the spectrogram of the auditory stimulus is projected into T tokens (one per time step) of embedding size E by means of a fully connected layer applied to each frequency vector. A learnable positional embedding is subsequently added to this compressed spectrogram representation before feeding it to a Transformer encoder with 1 layer, 4 heads, and a dimensionality of 48<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Vaswani, A. et al. Attention is all you need. In Proc. Advance Neural Information Processing Systems (eds Guyon, I. et al.) Vol. 30 &#010;                  https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf&#010;                  &#010;                . (Curran Associates, Inc., 2017).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR47\" id=\"ref-link-section-d61703220e2677\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a>. These hyperparameters were fixed for all datasets. As the outputs of the Transformer encoder are given as T processed tokens, we apply global average pooling over the token dimension and use the resulting tensor as the input to a final fully-connected readout layer, followed by a double-exponential activation function with per-unit learnable parameters. We observed empirically that the global average pooling operation is crucial to reach good performances while drastically reducing the size of the last fully connected layer.<\/p>\n<p>StateNet models<\/p>\n<p>A high-level schematic of the processing realized by our StateNet models is provided in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>c. Their architecture can be decomposed into three main elements: a downsampling layer, a stateful bottleneck, and a readout.<\/p>\n<p>Downsampling locally connected (LC) layer<\/p>\n<p>At each time step, we downsample the current vector of spectral information to reduce the dimensionality of input stimuli. Contrary to natural images, which are shift-invariant, spectrograms have very different statistics between low and high frequencies, thereby making weight sharing a less efficient computational strategy. Further motivated by the tonotopic organization observed along the auditory pathway<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 71\" title=\"Oliver, D. L. Ascending efferent projections of the superior olivary complex. Microsc. Res. Tech. 51, 355&#x2013;363 (2000).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR71\" id=\"ref-link-section-d61703220e2702\" rel=\"nofollow noopener\" target=\"_blank\">71<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 72\" title=\"Levy, R. B. &amp; Reyes, A. D. Spatial profile of excitatory and inhibitory synaptic connectivity in mouse primary auditory cortex. J. Neurosci. 32, 5609&#x2013;5619 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR72\" id=\"ref-link-section-d61703220e2705\" rel=\"nofollow noopener\" target=\"_blank\">72<\/a>, we use a LC layer with restricted receptive fields (as for convolutional layers) but with independent weights across frequency bands (see Supplementary Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">S2<\/a>). In other words, LC contains a subset of the weights of a fully connected (FC) layer, defined by a convolutional (CONV) connectivity pattern over the frequency axis. In theory, the performances obtained with this LC scheme are only a lower bound of what can be reached with FC. However, we found in practice that LC yields overall only slightly lower or similar results using the same hyperparameters (see\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary section<\/a> &#8220;Ablation study: connectivity in the first layer of StateNet\u201d), but with a smaller number of free learnable parameters, hence reducing the risks of overfitting and permitting a better generalization. LC also outperformed the CONV approach because it relaxes the weight-sharing constraint.<\/p>\n<p>Despite its biological and computational motivations, this approach has only rarely been incorporated into models of auditory processing. For example, Chen et al.<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 73\" title=\"Chen, Y.-h. et al. Locally-connected and convolutional neural networks for small footprint speaker recognition. In Proceedings of Interspeech 2015, 1136&#x2013;1140 (International Speech Communication Association, 2015).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR73\" id=\"ref-link-section-d61703220e2718\" rel=\"nofollow noopener\" target=\"_blank\">73<\/a> also used a local connectivity for speech recognition, but weight kernels were 2d (spectro-temporal) instead of 1d (only spectral and shared across the temporal dimension). In the field of computational neuroscience, Khatami and Escab\u00ed<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 74\" title=\"Khatami, F. &amp; Escab&#xED;, M. A. Spiking network optimized for word recognition in noise predicts auditory system hierarchy. PLoS Comput. Biol. 16, e1007558 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR74\" id=\"ref-link-section-d61703220e2722\" rel=\"nofollow noopener\" target=\"_blank\">74<\/a> imposed local Gaussian kernels as fully connected weights. Our implementation differs in that it implements the trade-off between CONV and FC, with fewer parameters than FC, and possibly faster execution.<\/p>\n<p>All in all, our proposed LC downsampling scheme is more biologically plausible than the FC and CONV alternatives, while providing a better trade-off between performances at the neural response fitting task and model complexity. In addition, it executes faster than the FC approach and prior LC implementations. An optimized PyTorch module is available on our code repository.<\/p>\n<p>Stateful bottleneck(s)<\/p>\n<p>It is composed of a single layer of either type of RNN, as adding more layers did not seem beneficial to performances in our preliminary experiments. RNNs are a type of artificial neural networks specifically developed to learn and process sequential inputs and notably temporal sequences. They work iteratively and build their output at each timestep from the current inputs as well as a constantly updated internal representation called hidden state. Because the mathematical details of the modules used here are fully provided in previous studies, we only report below their main properties. We invite the readers to consult the associated papers if a more thorough understanding of their computational principles is needed.<\/p>\n<p>Vanilla RNN<\/p>\n<p>In this paper, we designate &#8220;vanilla RNN&#8221; as the classical Elman network<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 48\" title=\"Elman, J. L. Finding structure in time. Cognit. Sci. 14, 179&#x2013;211 (1990).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR48\" id=\"ref-link-section-d61703220e2751\" rel=\"nofollow noopener\" target=\"_blank\">48<\/a> natively implemented in PyTorch, often considered as the most naive implementation of this class of models.<\/p>\n<p>Gated RNNs: LSTM and GRU<\/p>\n<p>A notorious problem with vanilla RNNs occurs when dealing with long sequences, as gradients can explode or vanish in the unrolled-over-time network<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 75\" title=\"Bengio, Y., Simard, P. &amp; Frasconi, P. Learning long-term dependencies with gradient descent is difficult. IEEE Trans. Neural Netw. 5, 157&#x2013;166 (1994).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR75\" id=\"ref-link-section-d61703220e2764\" rel=\"nofollow noopener\" target=\"_blank\">75<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 76\" title=\"Pascanu, R., Mikolov, T. &amp; Bengio, Y. On the difficulty of training recurrent neural networks. In Proc. 30th International Conference on International Conference on Machine Learning ICML&#x2019;13, Vol. 28, III-1310&#x2013;III-1318. arXiv:1211.5063.\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR76\" id=\"ref-link-section-d61703220e2767\" rel=\"nofollow noopener\" target=\"_blank\">76<\/a>, preventing them from exploiting long-range dependencies and therefore from performing well on large time scales. Gated RNNs such as LSTM<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 49\" title=\"Hochreiter, S. &amp; Schmidhuber, J. Long short-term memory. Neural Comput. 9, 1735&#x2013;1780 (1997).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR49\" id=\"ref-link-section-d61703220e2771\" rel=\"nofollow noopener\" target=\"_blank\">49<\/a> and GRU<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 50\" title=\"Cho, K. et al. Learning Phrase Representations using RNN Encoder&#x2013;Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1724&#x2013;1734 (Association for Computational Linguistics, Doha, Qatar, 2014).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR50\" id=\"ref-link-section-d61703220e2775\" rel=\"nofollow noopener\" target=\"_blank\">50<\/a> successfully circumvent these difficulties, and have imposed themselves as efficient modules for learning sequences in a recurrent approach.<\/p>\n<p>State-space models (SSMs): S4 and Mamba<\/p>\n<p>State Space Models (SSM) is a new class of models that takes inspiration from other fields of applied mathematics, such as signal processing or control theory<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 77\" title=\"Kalman, R. E. A new approach to linear filtering and prediction problems. Trans. ASME&#x2013;J. Basic Eng. 82, 35&#x2013;45 (1960).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR77\" id=\"ref-link-section-d61703220e2787\" rel=\"nofollow noopener\" target=\"_blank\">77<\/a>. Specifically designed for sequence-to-sequence modeling tasks (and thus for time series prediction as in the present study), they build upon the State-Space equations below with various parameterization techniques and numerical optimization methods.<\/p>\n<p>$$\\left\\{\\begin{array}{l}\\dot{x}(t)=Ax(t)+Bu(t)\\quad \\\\ y(t)=Cx(t)+Du(t)\\quad \\end{array}\\right.$$<\/p>\n<p>\n                    (2)\n                <\/p>\n<p>where \\(u\\in {\\mathbb{R}}\\) is the input, \\(x\\in {{\\mathbb{R}}}^{N}\\) the hidden state vector, \\(y\\in {\\mathbb{R}}\\) the output, and A, B, C and D system matrices.<\/p>\n<p>In particular, the original Structured State-Space Sequence (S4) model appears as one of the simplest versions of this paradigm, with only a few constraints on the system<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 51\" title=\"Gu, A., Goel, K. &amp; R&#xE9;, C. Efficiently Modeling Long Sequences with Structured State Spaces. In Proceedings of the International Conference on Learning Representations (ICLR, 2022).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR51\" id=\"ref-link-section-d61703220e3041\" rel=\"nofollow noopener\" target=\"_blank\">51<\/a>. At the opposite, the Mamba architecture is one of the more recent and sophisticated propositions in which system matrices are input-dependent, and has been shown to perform on par with Transformers on various benchmarks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 52\" title=\"Gu, A. &amp; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. In Proceedings of the Conference on Language Modeling (COLM, 2024). arXiv:2312.00752.\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR52\" id=\"ref-link-section-d61703220e3048\" rel=\"nofollow noopener\" target=\"_blank\">52<\/a>. As a whole, SSMs hold the promise of data-scalable models with great performances, while benefiting from a solid theoretical foundation, which permits to connect them to convolutional models (CNNs), RNNs with discrete timesteps, but also continuous linear time-invariant systems of ordinary differential equations. This last property is particularly interesting as it can ease the reverse-engineering process of a fitted model, thereby allowing for high levels of interpretability. In addition, trained SSMs can easily be modified to work at any temporal resolution, opening interesting use cases for the field of computational neuroscience and neural engineering.<\/p>\n<p>Readout<\/p>\n<p>The readout neural activity for a given unit is computed from a linear projection of the output state vector into a single scalar, repeated at each timestep.<\/p>\n<p>Electrophysiology datasets of auditory responses<\/p>\n<p>To characterize the ability of models to capture responses in the auditory cortex, we fitted them in a supervised manner on a wide gamut of natural audio stimulus-response datasets. These datasets were collected in different species (ferret, rat, zebra finch) and brain areas (MGB, AAF, A1, PEG, MLd, Field L), under varying behavioral conditions (awake, anesthetized) and using different recording modalities (spikes, membrane potentials). All of them are freely accessible on online repositories and were used with respect to their original license. Because the \u201cNS1&#8243;<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Harper, N. S. et al. Network receptive field modeling reveals extensive integration and multi-feature selectivity in auditory cortical neurons. PLoS Comput. Biol. 12, e1005113 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR22\" id=\"ref-link-section-d61703220e3073\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>, \u201cNAT4&#8243;<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Pennington, J. R. &amp; David, S. V. A convolutional neural network provides a generalizable model of natural sound coding by neural populations in auditory cortex. PLoS Comput. Biol. 19, e1011110 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR31\" id=\"ref-link-section-d61703220e3080\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 78\" title=\"Pennington, J. R. &amp; David, S. V. Can deep learning provide a generalizable model for dynamic sound encoding in auditory cortex? Neuroscience Preprint at &#010;                  https:\/\/doi.org\/10.1101\/2022.06.10.495698&#010;                  &#010;                 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR78\" id=\"ref-link-section-d61703220e3083\" rel=\"nofollow noopener\" target=\"_blank\">78<\/a>, and \u201cWehr&#8221;<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 24\" title=\"Machens, C. K., Wehr, M. S. &amp; Zador, A. M. Linearity of cortical receptive fields measured with natural sounds. J. Neurosci. 24, 1089&#x2013;1100 (2004).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR24\" id=\"ref-link-section-d61703220e3091\" rel=\"nofollow noopener\" target=\"_blank\">24<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 44\" title=\"Asari, C. K., Hiroki, M., Wehr, M. S. &amp; Zador, A. M. Auditory cortex and thalamic neuronal responses to various natural and synthetic sounds &#010;                  https:\/\/crcns.org\/data-sets\/ac\/ac-1\/about&#010;                  &#010;                . (2009).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR44\" id=\"ref-link-section-d61703220e3094\" rel=\"nofollow noopener\" target=\"_blank\">44<\/a> datasets have already been used in a previous study from our group and their pre-processing pipelines have been extensively described in the associated article<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Ran&#xE7;on, U., Masquelier, T. &amp; Cottereau, B. R. A general model unifying the adaptive, transient and sustained properties of on and off auditory neural responses. PLoS Comput. Biol. 20, 1&#x2013;32 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR27\" id=\"ref-link-section-d61703220e3098\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>, we only describe here the new datasets added to the current study. The corresponding data pre-processing methods are representative of what was performed in the previous datasets.<\/p>\n<p>AA1 datasets: MLd, field L (zebra finch)<\/p>\n<p>These two datasets consist in single-unit responses recorded from two auditory areas (MLd and Field L) in anesthetized male zebra finches, by Frederic Theunissen\u2019s group at UC Berkeley<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 43\" title=\"Theunissen, F. et al. Single-unit recordings from two auditory areas in male zebra finches. &#010;                  https:\/\/crcns.org\/data-sets\/aa\/aa-1\/about&#010;                  &#010;                 (2009).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR43\" id=\"ref-link-section-d61703220e3109\" rel=\"nofollow noopener\" target=\"_blank\">43<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 46\" title=\"Singh, N. C. &amp; Theunissen, F. E. Modulation spectra of natural sounds and ethological theories of auditory processing. J. Acoust. Soc. Am. 114, 3394&#x2013;3411 (2003).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR46\" id=\"ref-link-section-d61703220e3112\" rel=\"nofollow noopener\" target=\"_blank\">46<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 79\" title=\"Woolley, S., Fremouw, T., Hsu, A. &amp; Theunissen, F. Tuning for spectro-temporal modulations as a mechanism for auditory discrimination of natural sounds. Nat. Neurosci. 8, 1371&#x2013;9 (2005).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR79\" id=\"ref-link-section-d61703220e3115\" rel=\"nofollow noopener\" target=\"_blank\">79<\/a>. Stimuli were composed of short clips (&lt;5\u2009s) of conspecific songs to which animals had no prior exposure, and were modeled by log-compressed mel spectrograms with 32 frequencies ranging 0\u201316\u2009kHz, at a temporal resolution of 1\u2009ms. Extracellular recordings yielded a total of 50 single units in each area after spike sorting. Spike trains were binned in non-overlapping windows of 1 or 5\u2009ms, matching the resolution of the stimulus spectrogram. Sounds were presented with an average of 10 trials, and the PSTHs were obtained for each neuron and stimulus after averaging spike trains across trials and smoothing with a 21\u2009ms Hanning window<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 26\" title=\"Willmore, B. D., Schoppe, O., King, A. J., Schnupp, J. W. &amp; Harper, N. S. Incorporating midbrain adaptation to mean sound level improves models of auditory cortical processing. J. Neurosci. 36, 280&#x2013;289 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR26\" id=\"ref-link-section-d61703220e3119\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Hsu, A., Borst, A. &amp; Theunissen, F. Quantifying variability in neural responses and its application for the validation of model predictions. Network 15, 91&#x2013;109 (2004).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR80\" id=\"ref-link-section-d61703220e3122\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a>. For each neuron, recordings were performed in response to the same 20 audio stimuli, thereby allowing the training of a single model to predict the simultaneous activity of the whole population (i.e., a \u201cpopulation coding&#8221; paradigm).<\/p>\n<p>These data, also referred to as \u201cAA1&#8243;, are freely accessible from the CRCNS website (<a href=\"https:\/\/crcns.org\/data-sets\/aa\/aa-1\/about\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/crcns.org\/data-sets\/aa\/aa-1\/about<\/a>) and were used with respect to their original license.<\/p>\n<p>Asari dataset: A1, MGB (rat)<\/p>\n<p>This dataset consists in single-unit responses recorded from primary auditory cortex (A1) and medial geniculate body (MGB) neurons in anesthetized rats, by Anthony Zador\u2019s group at Cold Spring Harbor Laboratory<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 11\" title=\"Asari, H. &amp; Zador, A. M. Long-lasting context dependence constrains neural encoding models in rodent auditory cortex. J Neurophysiol. 102, 2638&#x2013;2656 (2009).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR11\" id=\"ref-link-section-d61703220e3150\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 44\" title=\"Asari, C. K., Hiroki, M., Wehr, M. S. &amp; Zador, A. M. Auditory cortex and thalamic neuronal responses to various natural and synthetic sounds &#010;                  https:\/\/crcns.org\/data-sets\/ac\/ac-1\/about&#010;                  &#010;                . (2009).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR44\" id=\"ref-link-section-d61703220e3153\" rel=\"nofollow noopener\" target=\"_blank\">44<\/a>. Stimuli were natural sounds typically lasting around 2\u20137\u2009s, and originally sampled at 44.1\u2009kHz and then resampled at 97\u2009kHz for presentation. Their associated cochleagrams were obtained using a short-term Fourier transform with 54 logarithmically distributed spectral bands from 0.1 to 45\u2009kHz, whose outputs were subsequently passed through a logarithmic compressive activation. The temporal resolution of stimulus cochleagrams and responses was set to 5\u2009ms. Because for each cell, recordings were performed in response to a different set of probe sounds, we could not apply the population coding paradigm, and one full model was fitted on the stimulus-response pairs for each unit. Contrary to the other datasets, recordings here are intracellular membrane potentials obtained through whole-cell patch-clamp techniques. One remarkable feature of these data is their very high trial-to-trial response reliability, making the noise-corrected normalized correlation coefficients of model predictions almost equal to the raw correlation coefficients (see \u201cPerformance metrics&#8221; subsection).<\/p>\n<p>Despite the good signal-to-noise ratio of this dataset, some trials are subject to recording artifacts and notably to drifts, which may be caused by motion of the animal and\/or of the recording electrode, or electromagnetic interferences with nearby devices. Note that drifts do not contaminate supra-threshold signals resulting from spike-sorted activity, such as PSTHs, because they are strictly positive and thus have a guaranteed stationarity. In order to remove these drifts, we detrended all responses using a custom approach further explicited in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary section<\/a> &#8220;Detrending with MedGauss filter\u201d.<\/p>\n<p>Similar to the previous dataset, these data can be found on CRCNS website (<a href=\"https:\/\/crcns.org\/data-sets\/ac\/ac-1\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/crcns.org\/data-sets\/ac\/ac-1<\/a>) as a subset of the \u201cAC1\u201d dataset.<\/p>\n<p>Neural response fitting task<\/p>\n<p>Because the current benchmark directly builds upon a previous study conducted by our group<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Ran&#xE7;on, U., Masquelier, T. &amp; Cottereau, B. R. A general model unifying the adaptive, transient and sustained properties of on and off auditory neural responses. PLoS Comput. Biol. 20, 1&#x2013;32 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR27\" id=\"ref-link-section-d61703220e3185\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>, performances and model trainings were conducted according to the same methods, which we describe again below.<\/p>\n<p>Task definition<\/p>\n<p>Neural response fitting is a sequence-to-sequence, time series regression task, taking a spectrogram representation \\(x\\in {{\\mathbb{R}}}^{F\\times T}\\) of a sound stimulus as an input, and outputting several 1d time series of neural response (one for each unit), \\(\\hat{r}\\in {{\\mathbb{R}}}^{N\\times T}\\). As the latter, we use the Peri-Stimulus Time Histogram (PSTH), which is the average recorded neural response across repeats. The loss function is the mean squared error (MSE) between the predicted time series and the recorded PSTH, and was evaluated for each time bin of each sequence:<\/p>\n<p>$${{{\\mathcal{L}}}}=\\frac{1}{NT}\\sum\\limits_{n=1}^{N}\\sum\\limits_{t=1}^{T}({\\hat{r}}_{n}[t]-{r}_{n}[t])$$<\/p>\n<p>\n                    (3)\n                <\/p>\n<p>where \\({\\hat{r}}_{n}[t]\\) is the predicted neural response for neuron n at time-step t, rn[t] the corresponding PSTH, N is the total number of recorded neurons to fit, and T is the total number of time-steps in the time series (to simplify the notations, we drop the time dependencies symbols [t] thereafter).<\/p>\n<p>Performance metrics<\/p>\n<p>The neural response fitting accuracy of the different models is estimated using the raw correlation coefficient (Pearsons\u2019 r), noted CCraw, between the model\u2019s predicted activity \\(\\hat{r}\\) and the ground-truth PSTH r, which is the response averaged over all M trials r(m):<\/p>\n<p>$$r=\\frac{1}{M}\\sum\\limits_{m=1}^{M}{r}^{(m)}$$<\/p>\n<p>\n                    (4)\n                <\/p>\n<p>$$C{C}_{raw}=\\frac{Cov(r,\\hat{r})}{\\sqrt{Var(r)Var(\\hat{r})}}$$<\/p>\n<p>\n                    (5)\n                <\/p>\n<p>where the covariance and variance are computed along the temporal dimension. Assuming neural variability is purely noise and given a limited number of stimulus presentations, perfect fits (i.e., CCraw\u2009=\u20091) are impossible to get in practice. In order to give an estimation of the best reachable performance given neuronal and experimental trial-to-trial variability, we use here the normalized correlation coefficient CCnorm, as defined in refs. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Hsu, A., Borst, A. &amp; Theunissen, F. Quantifying variability in neural responses and its application for the validation of model predictions. Network 15, 91&#x2013;109 (2004).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR80\" id=\"ref-link-section-d61703220e3791\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a>[,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 81\" title=\"Schoppe, O., Harper, N. S., Willmore, B. D. B., King, A. J. &amp; Schnupp, J. W. H. Measuring the performance of neural models. Front. Comput. Neurosci. 10, &#010;                  https:\/\/doi.org\/10.3389\/fncom.2016.00010&#010;                  &#010;                 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR81\" id=\"ref-link-section-d61703220e3796\" rel=\"nofollow noopener\" target=\"_blank\">81<\/a>. For a given optimization set (e.g., train, validation or test) composed of multiple clips of stimulus-response pairs, we first create a long sequence by temporally concatenating all clips together. We then evaluate the signal power SP in the recorded responses as:<\/p>\n<p>$$SP=\\frac{Var({\\sum }_{m = 1}^{M}{r}^{(m)})-\\mathop{\\sum }_{m = 1}^{M}Var({r}^{(m)})}{M(M-1)}$$<\/p>\n<p>\n                    (6)\n                <\/p>\n<p>which allows to compute the normalized correlation coefficient:<\/p>\n<p>$$C{C}_{norm}=\\frac{Cov(r,\\hat{r})}{\\sqrt{SP\\times Var(\\hat{r})}}$$<\/p>\n<p>\n                    (7)\n                <\/p>\n<p>When only one trial is available, we set CCraw\u2009=\u2009CCnorm, which corresponds to a fully repeatable recording uncontaminated by noise, thereby preventing any overestimation of performances by setting a lower bound in the absence of data.<\/p>\n<p>Optimization process\/model training<\/p>\n<p>All models were randomly initialized using default PyTorch methods and trained using gradient descent and backpropagation. Time recurrent models (DNet and StateNets) were trained using BPTT. Each training sample was a full stimulus-response pair whose duration varied between datasets because of the different recording protocols, but also within some datasets (AA1, Wehr, Asari) in order to maximize the amount of trial information for evaluation. We used AdamW optimizer<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Loshchilov, I. &amp; Hutter, F. Decoupled weight decay regularization. In Proc. 7th International Conference on Learning Representations, ICLR 6&#x2013;9, &#010;                  https:\/\/openreview.net\/forum?id=Bkg6RiCqY7&#010;                  &#010;                 (ICLR, 2019).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR82\" id=\"ref-link-section-d61703220e4129\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a> and its default PyTorch hyperparameters (\u03b21\u2009=\u20090.9, \u03b22\u2009=\u20090.999). We used a batch size of 1 for all datasets except NAT4, which have a limited number of training examples, and a batch size of 16 for both NAT4 datasets, which have consequently more (see &#8220;Supplementary Note <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>: Dataset and model details&#8221;). The learning rate was held constant during training and set to a value of 10\u22123. We found empirically that these values led to better results.<\/p>\n<p>We split each dataset into a training, a validation, and a test subset, respecting a 70\u201310\u201320% ratio as much as possible, depending on the number of stimulus-response pairs available for each cell in each dataset. After each training epoch, models were evaluated on the validation set, and if the validation loss had decreased with respect to the previous best model, the new model was saved. Models were trained until there was no improvement during 50 consecutive epochs on the validation set, at which point learning was stopped, the last best-performing model was saved, and evaluated on the test set. This procedure was repeated 10 times, each corresponding to a random seed, for different train-valid-test data splits and model parameters initializations, and the test metrics were averaged across splits. All models going through the exact same training pipeline (i.e., waveform-to spectrogram transform, training hyperparameters, etc.) ensured fair comparison between them, implying that architectures with higher test accuracy are genuinely better, despite potential differences with their original studies.<\/p>\n<p>Truncated backpropagation through time<\/p>\n<p>With regular BPTT, the entire RNN model is unrolled back in time, creating a graph that grows in size linearly with the sequence length. In addition, for most tasks, old time steps are less informative of the present than the most recent ones, and gradients can either vanish or explode (i.e., \u201cvanishing\/exploding gradient problem&#8221;)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 75\" title=\"Bengio, Y., Simard, P. &amp; Frasconi, P. Learning long-term dependencies with gradient descent is difficult. IEEE Trans. Neural Netw. 5, 157&#x2013;166 (1994).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR75\" id=\"ref-link-section-d61703220e4165\" rel=\"nofollow noopener\" target=\"_blank\">75<\/a>. TBPTT is a variation of this training algorithm aiming to alleviate these computational constraints<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 83\" title=\"Williams, R. J. &amp; Peng, J. An efficient gradient-based algorithm for on-line training of recurrent network trajectories. Neural Comput. 2, 490&#x2013;501 (1990).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR83\" id=\"ref-link-section-d61703220e4169\" rel=\"nofollow noopener\" target=\"_blank\">83<\/a>. Its principle relies on removing from the graph time steps older than a fixed temporal horizon K2. If the loss is evaluated every other K1 time steps, we note this algorithm TBPTT(K1,\u00a0K2). Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>. illustrates the underlying computations. In this study, in order to maximize the number of samples used for training, the loss is evaluated and backpropagated every time step, and therefore K1\u2009=\u20091. As a result, we simplify notations by referring to K2 as K. For the first t\u2009&lt;\u2009K2 time steps, the graph is built from the start of the sequence. For generality, the regular BPTT algorithm used to train our models in the main experiments corresponds to TBPTT(K1\u2009=\u20091,\u00a0K2\u2009=\u2009T), T being the sequence length. We also distinguish two sub-cases of TBPTT:<\/p>\n<p>TBPTT with warmup. The model is initialized with the default (null) hidden state h0\u00a0\u2009=\u20090 at the very start of the sequence (see Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>a). Inputs are then processed sequentially without building the computational graph, up to the K last time steps before loss evaluation. We qualify these first steps as a \u201cwarmup&#8221;. As a result, the graph starts with a model in an intermediary, non-default, and ecological hidden state resulting from this procedure.<\/p>\n<p>TBPTT without warmup. Here, warmup steps are skipped and the model is directly initialized to the default hidden state at the Kth time step before loss evaluation (see Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>b). Therefore, the model here does not have access to any prior information at all, making it a fairer comparison to the training of stateless models.<\/p>\n<p>                Fig. 8: Truncated backpropagation through time (TBPTT): methods.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s42003-025-08858-3\/figures\/8\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig8\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2025\/10\/42003_2025_8858_Fig8_HTML.png\" alt=\"figure 8\" loading=\"lazy\" width=\"685\" height=\"736\"\/><\/a><\/p>\n<p>a With warmup, the initial hidden state (purple) is the first of the training sequence. Subsequent time steps (in gray, delimited by \u201cno grad&#8221;) correspond to the warmup during which the hidden state is updated. In this example, loss evaluation is shown for two time steps (4 and 5), and their respective graph are colored in red and blue. K1 can be interpreted as the temporal stride between two loss evaluations and K2 (here, 3) as the maximum graph length, in number of time steps. b Without warmup, the first time steps are skipped and do not belong to the computational graphs. Therefore, the latter starts from the default, null hidden state instead of an intermediary value resulting from the warmup procedure above.<\/p>\n<p>Model interpretability with feature visualization: gradient maps and deep dreams<\/p>\n<p>We propose here a gradient-based iterative method which, for each unit of a trained neural network \\({{{\\mathcal{M}}}}\\), extracts its nonlinear receptive fields and estimates the auditory features xi that maximize its responses. This method builds on feature visualization techniques originally introduced in the AI community<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Wang, S. et al. Analysis of deep neural networks with the extended data Jacobian matrix. In Proc. 33rd International Conference on Machine Learning, ICML&#x2019;16, Vol. 48, 718&#x2013;727 (JMLR.org, 2016).\" href=\"#ref-CR38\" id=\"ref-link-section-d61703220e4350\">38<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Olah, C., Mordvintsev, A. &amp; Schubert, L. Feature visualization. Distill &#10;                  https:\/\/distill.pub\/2017\/feature-visualization&#10;                  &#10;                 (2017).\" href=\"#ref-CR39\" id=\"ref-link-section-d61703220e4350_1\">39<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 40\" title=\"Selvaraju, R. R. et al. Grad-cam: visual explanations from deep networks via gradient-based localization. In Proc. IEEE International Conference on Computer Vision (ICCV) 618&#x2013;626 (IEEE, 2017).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR40\" id=\"ref-link-section-d61703220e4353\" rel=\"nofollow noopener\" target=\"_blank\">40<\/a> and is known as gradient ascent, leveraging the fact that all the mathematical operations in our models are differentiable. The different steps of this approach are illustrated on Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>a and can be summarized as follows:<\/p>\n<p>                    1.<\/p>\n<p>As the first input to the model, use the null stimulus \\({x}_{0}\\in {{\\mathbb{R}}}^{F\\times T}\\), a uniform spectrogram of constant value (0 in our case). This initial stimulus is unbiased and bears no spectro-temporal information. From an information-theoretic perspective, it has no entropy. For a parallel with electrophysiology experiments, it is worth noting that the spectrogram of white noise is theoretically uniform too. The absence of spectro-temporal correlations within probe stimuli (which we respect here) is a strong theoretical requirement that led to the use of white noise in the initial development of the linear STRF theory<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Aertsen, A. &amp; Johannesma, P. The spectro-temporal receptive field-a functional characteristic of auditory neurons. Biol. Cybern. 42, 133&#x2013;43 (1981).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR17\" id=\"ref-link-section-d61703220e4419\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>, while further studies preferring natural stimuli used advanced techniques to correct their structure<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Theunissen, F. E., Sen, K. &amp; Doupe, A. J. Spectral-temporal receptive fields of nonlinear auditory neurons obtained using natural sounds. J. Neurosci. 20, 2315&#x2013;2331 (2000).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR18\" id=\"ref-link-section-d61703220e4423\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>.<\/p>\n<p>                    2.<\/p>\n<p>Pass x0 through the model and compute its outputs (i.e., the predicted time-series of activation for the whole neural population): \\(\\hat{r}={{{\\mathcal{M}}}}({x}_{0})\\in {{\\mathbb{R}}}^{N\\times T}\\)<\/p>\n<p>                    3.<\/p>\n<p>Define a loss \\({{{\\mathcal{L}}}}:{{\\mathbb{R}}}^{N\\times T}\\mapsto {\\mathbb{R}}\\) to minimize and compute it from the model prediction \\(\\hat{r}\\). In this paper, we only targeted a single unit n and used the opposite of its activation at the last (Tth) time-step: \\({{{\\mathcal{L}}}}(\\hat{r})=-{\\hat{r}}_{n}[T]\\). To compute the STRF associated with any neural population \\({{{\\mathcal{N}}}}\\), the following general loss can be used: \\({{{\\mathcal{L}}}}(\\hat{r})=-\\frac{1}{{{Card}(N)}}{\\sum}_{n\\in {{{\\mathcal{N}}}}}{\\hat{r}}_{n}[T]\\). The choice of this loss function permits to make the connection with the Spike-Triggered Average (STA) approach used by electrophysiologists, and where the stimulus instances preceding the discharge of the target unit are averaged<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Aertsen, A. &amp; Johannesma, P. The spectro-temporal receptive field-a functional characteristic of auditory neurons. Biol. Cybern. 42, 133&#x2013;43 (1981).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR17\" id=\"ref-link-section-d61703220e4846\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 41\" title=\"De Boer, E. &amp; Kuyper, P. Triggered correlation. IEEE Trans. Biomed. Eng. 15, 169&#x2013;179 (1968).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR41\" id=\"ref-link-section-d61703220e4849\" rel=\"nofollow noopener\" target=\"_blank\">41<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Schwartz, O., Pillow, J. W., Rust, N. C. &amp; Simoncelli, E. P. Spike-triggered neural characterization. J. Vis. 6, 13&#x2013;13 (2006).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR42\" id=\"ref-link-section-d61703220e4852\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a>. Models in our study were fitted to PSTHs or membrane potentials and thus output floating point values, which can be viewed as a spike probability. Trying to maximize this value at the present time-step (the last of the time series) by constructing previous stimulus time steps closely mimics STA. Maximizing the average firing rate across the whole stimulus presentation could be interesting to investigate in future studies.<\/p>\n<p>                    4.<\/p>\n<p>Back-propagate through the network the gradients of this loss \\({g}_{0}=\\frac{\\partial {{{\\mathcal{L}}}}}{\\partial {x}_{0}}\\), thereafter referred to as GradMaps. From their definition, these GradMaps can be directly related to linear STRF (see\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary text<\/a> &#8220;Bridging the gap between STRFs, gradMaps and dreams: theoretical framework\u201d).<\/p>\n<p>                    5.<\/p>\n<p>Use these gradients to perform a gradient ascent step and modify the input stimulus. For simplification, we denote this operation as x1\u2009=\u2009x0\u2009\u2212\u2009\u03b1g0, but some optimizers have momentum and more elaborate update rules. This is notably the case of Adam and AdamW<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Loshchilov, I. &amp; Hutter, F. Decoupled weight decay regularization. In Proc. 7th International Conference on Learning Representations, ICLR 6&#x2013;9, &#010;                  https:\/\/openreview.net\/forum?id=Bkg6RiCqY7&#010;                  &#010;                 (ICLR, 2019).\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#ref-CR82\" id=\"ref-link-section-d61703220e4956\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a>, which were used in our study. We did not use the SGD optimizer as it converged to much higher loss values\u2014so less optimal\u2014in preliminary experiments.<\/p>\n<p>                    6.<\/p>\n<p>Repeat these steps until an early stopping criterion is satisfied, in our case after a fixed number of iterations (1500, a value which led to sufficient loss decreases, see curves in Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>). The result of this process is an optimized input spectrogram d\u2009=\u2009xi that maximizes the activation of the target unit(s), thereafter referred to as Dream. If we define \\({{{\\mathcal{F}}}}:{{\\mathbb{R}}}^{F\\times T}\\mapsto {{\\mathbb{R}}}^{F\\times T}\\) as one iteration of the above process, such that \\({x}_{i}={{{\\mathcal{F}}}}({x}_{i-1}| {{{\\mathcal{M}}}},{{{\\mathcal{L}}}},n)\\) then recursively get \\({x}_{i}={{{\\mathcal{F}}}}\\circ {{{\\mathcal{F}}}}\\circ \\cdots \\circ {{{\\mathcal{F}}}}({x}_{0}| {{{\\mathcal{M}}}},{{{\\mathcal{L}}}},n)={{{{\\mathcal{F}}}}}^{i}({x}_{0}| {{{\\mathcal{M}}}},{{{\\mathcal{L}}}},n)\\).<\/p>\n<p>This approach does not assume any requirement on the model to interpret. It can produce infinitely long GradMaps, which are relatable to linear STRFs and can be implemented on any architecture, including RNNs like StateNet, but also stateless and transformers.<\/p>\n<p>Dream and GradMap energy<\/p>\n<p>A trace of temporal integration for a model can be simply defined from its GradMap g as the mean over all frequency bands f of squared elements for each latency t. We designate this measure the \u201cEnergy&#8221; of the GradMap:<\/p>\n<p>$$E[t]=\\frac{1}{F}\\mathop{\\sum }_{f=1}^{F}g{[f,t]}^{2}$$<\/p>\n<p>\n                    (8)\n                <\/p>\n<p>GradMap similarity matrix<\/p>\n<p>As shown in the corresponding results section, this matrix aims to identify functional clusters of models based on their GradMaps; it is built using the following methodology.<\/p>\n<p>The GradMaps of the models to compare are first computed with a number of time steps T slightly above their theoretical TRF size. In the case of the present study, stateless models had a TRF size of up to 43 time steps. Therefore, we computed GradMaps of T\u2009=\u200950 time steps for all models, including StateNet. The length of the GradMap should not be much greater than the theoretical TRF size of stateless models; otherwise, correlations between the GradMaps of the latter would artificially tend towards 1, as time steps beyond the receptive field do not receive any gradient and remain unaffected and at their initial value of 0. This step yields one GradMap per neuron, dataset, and model. After a flattening operation into a F\u2009\u00d7\u2009T vector, we compute Pearson\u2019s correlation coefficient as a pixel-wise metric to compare how similar GradMaps are between models, for a given neuron and dataset. The choice of the CC here instead of other distance metrics (e.g., Euclidean) is motivated by the fact that we are comparing the overall structure of the GradMap\/STRF (e.g., how inhibitory and excitatory regions are placed relative to each other) rather than its value. Furthermore, the goodness-of-fit of the Linear STRF model, to which we relate the GradMap, is insensitive to changes in scaling and shifting, precisely because the primary evaluation metric in the neural response modeling community is based on the CC too. As a result, if two models present the same GradMap up to an affine transformation, their functional similarity should be classified as perfect, which would not be the case with metrics, such as a pixel-wise MSE.<\/p>\n<p>Statistics and reproducibility<\/p>\n<p>One major contribution of this work is the sheer amount of implemented models and compiled datasets. All models were implemented and trained using PyTorch, a gold-standard library for deep learning in Python, leveraging autodiff. Datasets were pre-processed under the same format of a PyTorch Dataset class for convenience. Jobs required less than 2 GiB and were executed on Nvidia Titan V GPUs, taking tens of minutes to several hours, depending on the complexity of the model. As an example, the population training (5 seeds) of the StateNet GRU model trained in population coding on all 73 NS1 neurons in parallel typically takes less than 10\u2009min. Conversely, the single unit training (5 seeds) of the same model on the 21 neurons of the Wehr dataset takes more than 3\u2009h.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s42003-025-08858-3#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"We begin this section with a summary of common models of auditory neural responses and introduce a novel,&hellip;\n","protected":false},"author":2,"featured_media":66397,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[114,13866,3181,85,46,3183,48105,48106,48107],"class_list":["post-66396","post","type-post","status-publish","format-standard","has-post-thumbnail","category-business","tag-business","tag-cortex","tag-general","tag-il","tag-israel","tag-life-sciences","tag-network-models","tag-neural-encoding","tag-sensory-processing"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/66396","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=66396"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/66396\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/66397"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=66396"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=66396"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=66396"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}