Activation of prefrontal cortex and striatal regions in rats after shifting between rules in a T-maze

  1. Sidney I. Wiener1
  1. 1Center for Interdisciplinary Research in Biology (CIRB), College de France, Centre National de la Recherche Scientifique (CNRS), Institut National de la Santé et de la Recherche Médicale (INSERM), Université Paris Sciences et Lettres (PSL), Paris 75005, France
  2. 2Department of Child and Adolescent Psychiatry, New York University Medical School, New York, New York 10016, USA
  1. Corresponding author: sidney.wiener{at}college-de-france.fr
  1. 3 These authors contributed equally to this work.

Abstract

Prefrontal cortical and striatal areas have been identified by inactivation or lesion studies to be required for behavioral flexibility, including selecting and processing of different types of information. In order to identify these networks activated selectively during the acquisition of new reward contingency rules, rats were trained to discriminate orientations of bars presented in pseudorandom sequence on two video monitors positioned behind the goal sites on a T-maze with return arms. A second group already trained in the visual discrimination task learned to alternate left and right goal arm visits in the same maze while ignoring the visual cues still being presented. In each experimental group, once the rats reached criterion performance, the brains were prepared after a 90-min delay for later processing for c-fos immunohistochemistry. While both groups extinguished a prior strategy and acquired a new rule, they differed by the identity of the strategies and previous learning experience. Among the 28 forebrain areas examined, there were significant increases in the relative density of c-fos immunoreactive cell bodies after learning the second rule in the prefrontal cortex cingulate, the prelimbic and infralimbic areas, the dorsomedial striatum and the core of the nucleus accumbens, the ventral subiculum, and the central nucleus of the amygdala. These largely correspond to structures previously identified in inactivation studies, and their neurons fire synchronously during learning and strategy shifts. The data suggest that this dynamic network may underlie reward-based selection for action—a type of cognitive flexibility.

For survival, animals must adapt their behavior to changes in the external environment and to internal states. In goal-directed behavior, this can involve ignoring established stimulus reinforcement contingencies and shifting attention to a previously inconsequential stimulus. Previous work has shown that lesions of the medial frontal cortex lead to impairments in shifting between rules in rats (Joel et al. 1997; Ragozzino et al. 1999a,b). Birrell and Brown (2000) demonstrated that lesions of the prelimbic (PL) and infralimbic (IL) regions of the medial frontal cortex lead to impairment in shifting between strategies informed by different sensory modalities (extradimensional shifts), with no impairment in acquisition or reversal learning. Here, strategies are considered as patterns of responses with respect to cues and/or previous actions irrespective of rewards, while rules are strategies to be learned, since they regularly trigger rewards.

Elucidating the neural mechanisms underlying learning to select among alternative rules may involve different approaches, including simultaneous recording from multiple brain regions in animals performing cognitively demanding tasks (e.g., Bissonette and Roesch 2015; Oberto et al. 2022). However, recordings during behavior can only provide correlative evidence. Task-related neural activity can only be considered necessary for the execution of the task when the inactivation of the structure recorded is shown to impair performance. Nonetheless, the interpretation of inactivation studies is also limited, since the vicarious action of multiple overlapping and parallel networks in the brain can permit normal performance even when a vital structure is inactivated. Complementary to the latter approaches are techniques measuring the activation of brain areas. c-fos, an immediate early gene product, is a metabolic marker for neural activity (Dragunow and Faull 1989) and can be a powerful tool to identify brain regions and networks that are differentially activated for distinct behavioral contingencies (e.g., see Tronel and Sara 2002). Here, we identify multiple brain regions activated by the acquisition of a new rule using immunohistochemical marking of c-fos.

On the basis of previous lesion, inactivation, and neurophysiological recording studies, our analyses focused on brain areas associated with rule learning, including prefrontal cortices and their limbic striatal projection zones; limbic areas including the amygdala and the dorsal and ventral hippocampal subfields; and brainstem neuromodulatory centers (e.g., Birrell and Brown 2000; Burnham et al. 2010). We also examined dorsal striatal zones implicated in motor and habit learning, which were not expected to be activated during learning a new rule. A particular motivation for the present work is the observation that once rats learned and thus reliably received rewards for performing a behavioral strategy in goal-directed behavioral tasks, prefrontal cortical pyramidal neurons as well as ventral striatal neurons shifted their firing phase relative to hippocampal theta oscillatory activities (Y-maze [Benchenane et al. 2010] and T-maze [Oberto et al. 2022]), leading to the formation of synchronous cell assemblies. Furthermore, after learning, prefrontal interneurons more effectively inhibited principal neurons, and synchronously active prefrontal neuronal assemblies appeared for the first time (Peyrache et al. 2009). Synchronous cell activity is associated with c-fos expression elsewhere in the brain (Guo et al. 2007). While these studies have focused on parts of the prefrontal cortex and striatum, the present study examines c-fos immunoreactivity to identify the extent of the network reorganization related to learning of a new behavior–reward contingency and suggests associated areas to examine for synchrony during new rule learning.

Results

Behavioral performance

Rats were tested in cohorts of two, with one achieving criterion performance in a visual orientation discrimination (VOD) task, and the other continuing on to learn alternation (the VOD-ALT group) in the same maze with the same cues. The VOD group reached performance in 17.8 ± 4.5 (SEM) sessions, comprising 759 ± 240 trials over four to 12 training days. The VOD-ALT group reached the VOD performance criterion in 13.2 ± 4.7 sessions comprising 634.4 ± 210.4 trials. The number of trials and days to VOD criterion in the VOD and VOD-ALT groups was not significantly different (two-way paired Student's t-test, P = 0.53 for trials and P = 0.30 for days). Training for criterion performance in the ALT task required 4.2 ± 0.36 sessions comprised of 311 ± 20.5 trials (see Fig. 1D).

Figure 1.

(A) As the rat crossed the middle of the central arm of the T-maze with return arms, cues appeared on the two screens. Correct choices were rewarded for going to the arm with vertical stripes in the VOD task and for going to the opposite arm from the previous trial in the ALT task. (B) Time course of experimental protocol for each cohort of rats. (C,D) Learning curves for the respective rats for the acquisition of the VOD task (C) and the ALT task (D). (E) Representative histological sections showing immunoreactive neurons. (F) Sites sampled for neuron counts (adapted from Paxinos and Watson 1998). (G) Relative density of c-fos immunoreactivity. (*) P < 0.05 significance in a one-tailed paired t-test for the five cohorts of rats. (Cg1) Cingulate area 1, (PL) prelimbic area, (IL) infralimbic area, (VO) ventral orbital cortex, (LO) lateral orbital cortex, (D) dorsal striatum, (DL) dorsolateral striatum, (DM) dorsomedial striatum, (NAcC) nucleus accumbens core, (NAcS) nucleus accumbens shell.

During training in both tasks, the rats spontaneously used other strategies (such as go left, go right, ALT during the VOD task, and VOD during the ALT task) despite the fact that this led to rewards on only 50% of the trials, on average. The principal tendency was for spontaneous alternation during the VOD training. To evaluate this tendency, only runs of six or more successive visits to the arm opposite the one visited on the previous trial were counted (binomial probability P ≤ 0.0325), and these constituted 30% ± 7% of the trials during training in the VOD group.

As expected, after the VOD acquisition and the switch to the ALT rule, the rats all initially persisted in the VOD strategy. Extinction occurred at least two sessions before they reached criterion performance in the ALT rule. In all cases, they once again performed a minimum of six successive VOD trials in the penultimate or final session prior to reaching the ALT criterion. This indicates that strategy shifting had indeed taken place when the ALT rule was acquired, since the most recent performance of the VOD strategy was in the recent past.

c-fos labeling

Tissues from each cohort pair were processed together in order to permit comparisons between members of the respective pairs to be unaffected by variations in c-fos immunohistochemical processing (Fig. 1E). The VOD-ALT group had significantly higher levels of c-fos immunoreactivity relative to the VOD group in the cingulate (Cg1), prelimbic (PL), and infralimbic (IL) cortices (t(4) = 2.58, t(4) = 2.40, and t(4) = 2.48, respectively; pairwise Student's t-test, one-tailed, P < 0.05) (Fig. 1G, top left). See Figure 1F for the anatomical localization of the regions sampled. No significant increase was detected in ventral orbital (VO) or lateral orbital (LO) areas. In the striatum (Fig. 1G, top right), significantly higher levels of c-fos immunoreactivity were observed in limbic zones: the dorsomedial (DM) striatum and the core of the nucleus accumbens (NAcC) (t(4) = 2.67 and t(4) = 2.19, respectively; pairwise Student's t-test, one-tailed, P < 0.05) (cf. Fig. 1F). However, no such increases appeared in the dorsal (D) or dorsolateral (DL) striatum or in the nucleus accumbens shell (NAcS). Data from the individual rats are shown in Figure 2 and suggest nonsignificant tendencies for increases in VO and NAcS (pairwise Student's t-tests, P = 0.074 and P = 0.11). Note that the dorsal (D) striatum measurements were limited to the dorso–central striatum (corresponding to the sensorimotor striatum) and did not extend medially to the limbic zone occupied by the DM striatum.

Figure 2.

Individual data for each cohort as shown in Figure 1G, with asterisks indicating the significant increases in Figure 1G. In the latter areas, the majority of the cohorts also show trends for increases. Abbreviations are the same as in Figure 1G.

Burnham et al. (2010) observed that c-fos immunoreactivity levels in the orbitofrontal cortex are closely related to the number of trials performed on the final training day. Here, the number of trials in the VOD and ALT groups on the final testing day did not show a significant difference (two-tailed paired t-test, P = 0.22), and thus this is unlikely to be a possible confound.

In the hippocampal formation (Fig. 1G, bottom), there was a rather low density of c-fos immunoreactive neurons with no significant differences between the two groups in the dorsal or ventral CA1, CA3, or dentate gyrus. Despite the small number of marked neurons, ventral subiculum c-fos immunoreactivity was significantly higher in the VOD-ALT group (t(4) = 3.39; pairwise Student's t-test, one-tailed, P < 0.05), while there was no difference between the groups in the dorsal subiculum. A significantly higher level of c-fos for the VOD-ALT group also appeared in the central nucleus of the amygdala (t(4) = 2.30; pairwise Student's t-test, one-tailed, P < 0.05) (data not shown), while the difference did not reach significance in the basolateral nucleus of the amygdala.

Note that although not statistically significant, the mean values of immunoreactive relative density tended to be higher in the VOD-ALT group than in the VOD group in all but one of these hippocampal system structures (the exception was ventral CA3, which also had extremely low densities). Brainstem neuromodulatory centers were also examined (substantia nigra pars compacta and reticulata, dorsal raphé nucleus, laterodorsal tegmental nucleus, ventral tegmental area, raphé nucleus, dorsal tegmental nucleus, and locus coeruleus), and all had very low relative densities of marked neurons. In those with sufficient staining to permit analyses, there were no significant differences between the two groups (data not shown).

Discussion

The principal findings here are that the density of c-fos immunoreactive neurons increases in the prefrontal cingulate, prelimbic, and infralimbic areas; the dorsomedial striatum; the core of the nucleus accumbens; and the ventral subiculum and central nucleus of the amygdala in animals after they learned a new rule (ALT) compared with animals having learned the VOD rule.

Extinction of the VOD rule occurred at least two sessions before rats achieved criterion performance in the ALT task, and thus any c-fos activity changes related to extinction would have subsided by then (Bertaina-Anglade et al. 2000). It is important to note that animals had spontaneously performed the ALT strategy during VOD training; therefore, the network underlying this strategy had been recently active, and thus this alone would not be expected to increase c-fos activity in our samples (Chung 2015). Rather, the association of this strategy with consistent reward (that is, the acquisition of the rule) is more likely to be responsible, with c-fos activity specifically associated with this stage of learning and not previous ones (as observed in a different task in mice by Bertaina-Anglade et al. 2000). The increase in c-fos in the central nucleus (CN) of the amygdala and the ventral subiculum (which project to the nucleus accumbens) (Aylward and Totterdell 1993) may be related to the change in the reward value associated with maze cues upon acquisition of the ALT task. The nucleus accumbens activity is also consistent with its afferents from the prefrontal cortical areas with increased c-fos, as well as their mutual participation in highly synchronous neuronal assembly activity during shifting between comparable tasks in this maze (Oberto et al. 2022).

Relations to neurophysiological recording studies

There are anatomical connections among these zones (Jay and Witter 1991; Fudge et al. 2002; Mailly et al. 2013), and thus they are likely to compose functional networks. The absence of significant differences in the dorsal hippocampus (and associated dorsal subiculum) is consistent with previous neurophysiological recordings of the hippocampus and nucleus accumbens of rats in a plus maze as they alternated between spatial orientation strategies using distal cue configurations versus visible beacons. In that study, we observed that hippocampal neurons do not change their firing fields between the two tasks (Trullier et al. 1999) but that nucleus accumbens neurons do change their firing patterns after task changes and platform rotations (Shibata et al. 2001). Thus, in these studies, the rule shifting was reflected by changes in accumbens neural activity but, in Trullier et al. (1999), not in hippocampal activity.

The present c-fos results are consistent with our previous electrophysiological observations of changes in circuit properties in the PL upon new learning of a rewarded rule: Hippocampal inputs are more effectively transmitted by inhibitory interneurons to principal neurons, and the latter change their phase of firing relative to hippocampal theta oscillations. Synchronously active cell assemblies are then formed, likely in relation to dopaminergic neuromodulatory inputs (Benchenane et al. 2010). Thus, the reward expectation that comes with actual learning of the behavior–reward contingency would be expected to bring about changes in the network underlying the behavior, and this would be reflected in the c-fos activity. Furthermore, in rats shifting between visual and spatial discrimination rules in this T-maze, we found changes in the activation of prefrontal and accumbens neurons and their cell assemblies (Oberto et al. 2022).

Relation to inactivation studies

A popular experimental paradigm for testing task switching is to train rats in “place” and “response” rules in a plus maze. The place task rewards visits to a specific arm, while the response strategy requires the same body turn at the choice point. Rich and Shapiro (2007) inactivated the PL/IL areas with muscimol injections and, in contrast to the present results, observed that task switch acquisition was not impaired, but memory for the recently acquired switch was indeed impaired 24 h later. Instead, the rats persevered at the initially learned strategy. However, Ragozzino et al. (1999a) did observe switching impairments with reversible inactivation of the PL/IL areas by local tetracaine infusion. Similar impairments in rule shifting were observed by Oualian and Gisquet-Verrier (2010) after PL and/or IL infusions with ibotenic acid. A possible explanation for this discrepancy advanced by Rich and Shapiro (2007) would be that alternate networks would support this acquisition. Our data would suggest that Cg1 could be part of such a network.

Relation to other immediate early gene studies

Burnham et al. (2010) examined c-fos immunoreactivity in rats acquiring an extradimensional shift between rules in a digging task with odor and texture cues. They found increased activity relative to cage controls in the medial prefrontal cortex and orbitofrontal cortex. In a study of male F-344 rats shifting between response and place strategies in the plus maze, Grella et al. (2013) observed more activation of the immediate early gene arc in the PL and IL areas and VO cortex (but not MO and LO cortices) than in cage control animals. While the tendency for increased c-fos activity in the VO cortex in VOD-ALT animals did not reach significance here, this does not preclude its participation in rule shifts. The use of cagemate controls in the latter studies did not control for the increased sensory and motor activity of the maze training.

The present results show, for the first time, striatal activation during the acquisition of a new rule requiring selection among orienting cues and point to a prefrontal striatal network underlying this behavioral and cognitive flexibility, consistent with lesion and pharmacological manipulation studies (Bissonette and Roesch 2017). While this immediate early gene product study identifies members of this network, electrophysiological studies will be required to determine with temporal precision the dynamic roles of these brain regions in switching attention from one stimulus modality to another and in modifying behavior according to new reward contingencies. Further studies could also apply reward contingencies with other discriminative stimuli in order to determine whether this network is involved in a variety of rule shifts. Since c-fos activation can reflect synchronous cell activity (Guo et al. 2007), one prediction is that during rule shift behaviors, cell assembly activations could emerge in a broad network, including prefrontal, limbic striatal, ventral subicular, and amygdalar neurons.

Materials and Methods

Ten male Long-Evans rats (250–300 g; Janvier) were housed singly in cages under a 12-h light–dark cycle. All rats were weighed and handled daily after their arrival. They had free access to water at first, and then were restricted to 10 min per day (or more as required) to no less than 85% of their free-feeding weight to motivate them to search for liquid rewards during pretraining and training days.

The T-maze (Fig. 1A) was constructed from wood and painted matte black. The central stem and top alley were 1 m long and 8 cm wide with 2-cm-high borders. The maze was elevated 70 cm above the floor and surrounded by a black cylindrical curtain 3 m in diameter running from floor to ceiling. Visual cues were displayed on two TV screens (80 cm diagonal) located behind the reward arms. When the rat crossed a photodetector beam on the central arm, vertical bars were displayed on one and horizontal bars were displayed on the other in a pseudorandom order (spatial frequency of 0.13 cycles per degree, viewed from the trigger point). Each of the 16 possible left–right configurations of visual cues was displayed at equal incidences over consecutive series of four trials, with the exception of three repetitions of the same cue to the left or the right (thus, 14 sequences were used). Following correct choices, 30 µL of 0.25% saccharinated water rewards was dispensed from small wells operated by solenoid valves controlled by a CED Power 1401 system (Cambridge Electronic Design).

The rats were first habituated to the T-maze over 3–5 d by allowing them to forage for Kellogg's Coco-Pops (which they had previously sampled in their home cages) that were distributed randomly along the surface. Habituation was terminated after the rats moved about easily and consumed food on all parts of the maze without freezing or defecating. Rats were then pretrained over 3–9 d to run in the correct direction in the maze, starting from the base of the central arm, then turning left or right at the “T,” and proceeding to the reward site. At first, guillotine doors (controlled remotely by a manual pulley system) prevented backtracking after entries into the selected reward arm. During pretraining, a drop of saccharinated water was delivered for either choice. The rat then had to continue back along the return arm located on the same side to start a new trial, and a movable barrier prevented direction reversal. If the rat turned in only one of the directions at the choice point, a barrier was put into place for several trials to force visits to the other side. Pretraining ended after the rats became proficient at performing, moving in the correct direction without guidance.

The experiment was carried out in five replications, each with a cohort of two rats trained in successive sessions, and their brain tissues were processed together (Fig. 1B). This permitted reduction of variability and hence a reduced sample size, since all comparisons involved only pairs of rats exposed to identical housing and training conditions, and tissues were processed under identical conditions with identical reagents. They were first trained in a VOD task. One rat of the cohort was then anesthetized and perfused while the other was trained in an alternation task (ALT) in the same T-maze. Rats were selected randomly for participation in the two respective groups. In the VOD task, the screen with the vertically oriented black stripes was behind the arm associated with reward on that trial, while in the ALT task rats were rewarded for choosing the reward arm opposite to the one selected on the previous (rewarded or unrewarded) trial, ignoring the visual cues. For unrewarded choices, the rats had to continue down the return arm to initiate a new trial. The first trial of ALT sessions was not rewarded. The onset of the ALT task was signaled by an intermittent tone (repetition of the Microsoft Windows standard system sound “Asterisk”) presented from a loudspeaker in front of the T-maze simultaneously with the visual cues. The tone continued until the rat completed a turn onto a reward arm in each trial. Otherwise, all stimuli remained the same as in the VOD task. The VOD-ALT group first learned the VOD task and then was trained to criterion in the ALT task. Note that this experimental design emulates those used in brain imaging studies. In order to model the neurophysiological recording protocol of Peyrache et al. (2009) and Benchenane et al. (2010) and to accelerate the training process, rats were permitted as many trials as they would run in a given session, and up to three sessions were run per day. Sessions were stopped when the animal stopped moving, and all sessions lasted <30 min. Intervals between sessions were at least 3 h.

To establish the learning criterion for the VOD task, we counted the number of runs of four or more successive correct trials, corresponding to a binomial probability of P ≤ (0.5)4 or P ≤ 0.0625. The total number of trials in these runs had to exceed 50% over two consecutive days. Alternatively, a total of 80% correct trials (irrespective of run length) in a single session was also considered to satisfy the learning criterion (see Fig. 1C). The rat was perfused 90 min later, since this was reported to be the optimal delay for revealing activation-dependent c-fos accumulation (Sheng and Greenberg 1990). The criterion of VOD performance for VOD-ALT group prior to ALT training was 80% correct (irrespective of run length) over 50 successive trials (which could be from the last two sessions, if the final one was too brief). The learning criterion for the ALT task was a run of >12 alternation trials in a given session, and again a 90-min delay preceded perfusion. These elevated criteria were necessary to assure reliable task performance.

Rats were administered a lethal dose of pentobarbital and perfused with saline followed by freshly prepared 4% paraformaldehyde in phosphate buffer (PB; pH 7.4). Brains were removed, postfixed in the buffered 4% paraformaldehyde solution for 24 h, and then placed in PB containing 30% sucrose (pH 7.4). Frozen sections were cut coronally at 40 µm. Sections from brains of the two rats belonging to the same cohort (one VOD rat and one VOD-ALT rat) were processed simultaneously. Sections were first incubated in 0.3% H2O2 for 30 min. After three rinses, the sections were incubated overnight at room temperature in PB containing 0.2% primary c-fos antibody (Santa Cruz Biotechnology sc-52), 0.5% Triton X-100, and 0.1% bovine serum albumin (BSA; Sigma). After three rinses, the sections were incubated during 2 h in PB containing 0.5% secondary biotinylated antibody, 0.5% Triton X-100, and 0.1% BSA. After four rinses, the sections were incubated in avidin–biotin–peroxidase complex (ABC; Vector Laboratories) solution for 2 h. Finally, c-fos immunoreactivity was revealed with 0.05% diaminobenzidine (DAB).

Regions of interest were carefully delineated and verified on adjacent sections treated with Nissl stain. c-fos immunostained neurons were counted automatically with Fiji image processing software (http://fiji.sc/Fiji) by two observers blind to treatment groups. Paired t-tests were used for statistical evaluation, with P < 0.05, one-tailed, as the level of significance. One-tailed tests were used because the goal of the study was to determine the extent of the network with increased activity after the new acquisition relative to the initial acquisition-only increases that were expected. The statistics were not corrected for multiple comparisons, since all tests were independent and planned: Most regions were selected for measurements based on earlier anatomical, inactivation, lesion, or recording studies where activity was associated with rule acquisition (while dorsal and DL striatum were tested as control areas unlikely to be activated because of their sensorimotor functions).

All experiments were approved by the local animal experimentation ethics committee (Comité d'Ethique en Matière d'Expérimentation Animale-Paris Centre et Sud 59) and were in accord with international standards (Directive 86/609/EEC; ESF-EMRC position paper 2010/63/EU; National Institutes of Health guidelines) and legal regulations (certificate no. 7186, Ministère de l'Agriculture et de la Pêche) regarding the use and care of laboratory animals.

Competing interest statement

The authors declare no competing interests.

Acknowledgments

We thank France Maloumian for expert help preparing figures, and Marie-Annick Thomas and Suzette Doutremer for training and advice for histology. We thank Nicole Quenech'du and Jérémie Teillon for advice and training for image processing, and Dr. Anne Cei and Dr. Michaël Zugaro for comments on the manuscript. Support came from French National Research Agency ANR-2010-BLAN-0217-01 Neurobot. A.B. was supported by the Erasmus Program of the European Community. We thank MemoLife Laboratory of Excellence and Fondation Bettencourt Schueller.

Author contributions: S.I.W., H.G., and V.O. developed the experimental design. S.J.S. supervised the initial experiments. H.G., V.O., and A.B. performed the experiments. All authors analyzed the data. S.I.W., H.G., V.O., and S.J.S. wrote the manuscript. All authors approved the final version of the manuscript.

  • Received May 23, 2023.
  • Accepted June 21, 2023.

This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the first 12 months after the full-issue publication date (see http://learnmem.cshlp.org/site/misc/terms.xhtml). After 12 months, it is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.

References

| Table of Contents