The role of the parahippocampal cortex in memory consolidation for scenes

  1. Kareem A. Zaghloul2
  1. 1Department of Psychology, University of Maryland, College Park, Maryland 20742, USA
  2. 2Surgical Neurology Branch, NINDS, National Institutes of Health, Bethesda, Maryland 20892, USA
  3. 3Department of Neurosurgery, University of Maryland, Baltimore, Maryland 21201, USA
  4. 4Laboratory of Brain and Cognition, NIMH, National Institutes of Health, Bethesda, Maryland 20892, USA
  1. Corresponding authors: zanexie{at}umd.edu; bakerchris{at}mail.nih.gov; kareem.zaghloul{at}nih.gov

Abstract

Classic models propose that forming lasting visual memories involves coordinated interactions between visually selective neocortical structures and the hippocampus during memory consolidation. However, the precise role of visually selective neocortical structures in memory consolidation remains elusive, given their potential contributions spanning from initial perceptual encoding to subsequent memory reactivation. We capitalized on a unique opportunity, involving direct recording from the posterior parahippocampus and its subsequent resection in a neurological patient, to investigate the impact of scene-selective neocortical lesions on visual memory consolidation. First, with intracranial EEG, we confirmed the functional relevance of the patient's resected tissues in representing a specific visual category, in this case, scene images. Subsequently, we identified disruption of memory for scenes relative to faces and objects during the participant's postoperative visit. This finding prompted a comprehensive analysis of visual memory across different visual categories in this participant, as well as an examination of similar functions in other neurological patients with intact parahippocampi and a cohort of online participants. Through these within- and between-participant comparisons, we identified greater time-dependent reduction in visual memory for scene images following the resection of the posterior parahippocampus. Importantly, these changes in memory retention could not be attributed to a general reduction in initial memory encoding following neocortical lesions. Our findings, therefore, suggest that reactivating scene-selective neocortical areas is essential for converting transient visual perceptual experiences into lasting long-term scene memories.

Formation of enduring memories requires a consolidation process to transform transient perceptual experiences into durable memory representations (McGaugh 2000). In the visual domain, this process is often hypothesized to involve coordinated interactions between the hippocampus and neocortical structures along the ventral pathway (McIntyre et al. 2012; Collins and Dickerson 2019). Following initial encoding, the hippocampus is postulated to guide the reorganization of neocortical information by reactivating initially engaged regions to establish lasting memories spanning hours, days, or even years (Squire et al. 2015). Lesions in the hippocampus can thus lead to an overall memory deficit, especially after a long delay (Nadel and Moscovitch 1997; Squire et al. 2015). Complementing this hippocampal involvement (McClelland et al. 1995), neocortical structures are believed to support memory consolidation by aiding in the retention of information encoded in these regions (Wiltgen et al. 2004). As such, reactivating neocortical regions specialized in perceiving specific stimuli (e.g., faces and scenes) would facilitate memories of these stimuli over time (Polyn et al. 2005). Yet, as memories usually undergo transformations during consolidation and may be represented differently compared to their initial perceptual representations (Winocur and Moscovitch 2011; Dudai et al. 2015), the extent to which category-selective neocortical regions are required for the consolidation of different types of memories has remained unclear.

Studies of neurological patients with focal neocortical damage beyond the hippocampus present an opportunity to clarify this issue (De Renzi 1982; Vaidya et al. 2019). Past research suggested that neocortical lesions may lead to specific memory deficits in contrast to the overall memory deficits often observed in cases of hippocampal lesions (Milner 1968; De Renzi 1982). For instance, patients who have undergone anterior temporal lobectomy often retain visual memories despite some verbal challenges (Milner 1968). However, the interpretation of existing findings in many of these patients is complicated by various factors, including brain plasticity (Kolb and Whishaw 1989; Holtmaat and Svoboda 2009) and compensation from other intact brain areas (Celeghin et al. 2017). Moreover, the limited knowledge about the resected brain areas prior to the lesion in most previous case-control studies further adds to this complexity.

Leveraging the unique opportunity to investigate the effects of focal resections in intractable epilepsy, we conducted a longitudinal study in a participant (VH) who underwent targeted surgical resection of the posterior temporal regions for seizure treatment (Fig. 1A). The surgical resection involved the conventional left parahippocampus place area (PPA) (Epstein and Kanwisher 1998), while sparing the well-recognized fusiform face area (FFA) (Kanwisher et al. 1997), confirmed by pre-lesion intracranial EEG (iEEG) and post-lesion fMRI localization (Fig. 1B,C). We were interested in investigating two possibilities that could result from such a selective neocortical lesion. First, this lesion could result in selective memory deficits for visually encoded stimuli related to the PPA, in this case scene images, considering that ventral neocortical lesions could selectively impair perceptual encoding for distinct visual categories (Habib and Sirigu 1987; Barton et al. 2002; Mendez and Cherrier 2003). This would manifest as a consistent reduction in overall memory likelihood for scenes across short and long delays (Fig. 1D). Alternatively, if perception and memory for scenes recruit separate functional networks (Steel et al. 2021, 2023) and compensatory mechanisms are recruited to facilitate initial perceptual encoding (Gramsbergen 2007), this lesion might spare initial encoding while interrupting the reactivation of scene information required for memory consolidation. This would result in scene-specific deficits that only emerge over time (Fig. 1E; Ploner et al. 2000). Clarifying these possibilities is key to elucidating how visual information is transferred into durable memories—a fundamental process underlying visual cognition (Chun and Johnson 2011).

Figure 1.

Design and predictions for the current case-control study. (A) Patient VH underwent a surgical trial for epilepsy, involving intracranial electrode implantation for seizure monitoring. VH completed a visual task during monitoring (see Fig. 2 for details). Following clinical evaluation, surgery removed posterior medial temporal lobe tissue. VH later participated in memory tasks at the 12-month (see Fig. 3) and 36-month (see Fig. 4) follow-ups using diverse visual stimuli. We tested 165 controls to assess VH's memory performance compared to other neurological patients or online participants. (B) VH's lesion (marked in pink) overlaps with the conventional PPA (marked in yellow) identified based on a probabilistic functional atlas from the previous research (Rosenke et al. 2021). In contrast, VH's lesion barely overlaps with the conventional FFA in the more posterior side (marked in teal). All regional masks are normalized to the MNI space for visualization and comparison. (C) Selectivity for scenes and faces in VH's brain following resection at 36-month follow-up revealed by fMRI. t-maps resulting from the generalized linear model contrast faces > scenes are overlaid on the anatomical volume and thresholded at t ± 4.0. Note that selectivity for scenes extends throughout the anterior parahippocampal gyrus in the intact right hemisphere, but the corresponding anterior portion is absent in the left hemisphere due to the lesion. (D) VH's lesions in PPA may lead to a generalized deficit in processing scene images during initial perception, leading to an overall reduction in subsequent memory performance. (E) Alternatively, compensatory mechanisms from other brain areas might facilitate initial encoding, while PPA damage primarily disrupts the reactivation of specific neocortical regions essential for long-term consolidation of scene memories. As such, emerging deficits might become more pronounced during the consolidation phase, potentially leaving initial encoding minimally affected.

Results

We performed a targeted neocortical resection in a 36-year-old right-handed male (Patient VH) with drug-resistant epilepsy for treatment of his seizures (see Case history). Prior to his surgical resection, VH underwent a monitoring procedure involving intracranial implantation of platinum stereotactic EEG recording contacts to localize epileptogenic brain regions. Clinical assessment based on the iEEG recordings suggested a potential seizure onset zone near the posterior medial temporal lobe. Following the monitoring period, VH therefore underwent a surgical procedure that involved resecting parts of the left parahippocampus, its adjacent tissues in the inferior temporal gyrus and the fusiform gyrus, along with a portion of the occipitotemporal cortex (see Fig. 1B).

Following his initial operative procedure, we invited VH for postsurgical clinical evaluation, neuroimaging, and/or behavioral testing at 12- and 36-month intervals (Fig. 1A). To gauge VH's memory task performance relative to other patients and controls, we enrolled 20 neurological patients with intact bilateral parahippocampi to perform the same memory tasks that VH completed (10 females; 37.15 ± 2.04 [mean ± s.e.m.] years old; Wechsler intelligence quotient (IQ): 85.06 ± 3.82; see Supplemental Table S1). We also recruited an additional control group of 145 online participants via Amazon Mechanical Turk for the immediate and delayed recognition memory task (35.64 ± 0.78 years; 65 males).

Resected brain tissues in VH overlapped with the PPA

We first characterized VH's lesion using available structural and functional data (Fig. 1B). We first aligned VH's resected brain regions to his preoperative MRI and then normalized to a standard space (MNI), where probability atlases of the conventional FFA and PPA from healthy observers have been previously defined (Rosenke et al. 2021). Notably, the resected area exhibits a stronger overlap with the conventional PPA (86.5% overlap), when compared with the conventional posterior FFA on the left hemisphere (0% overlap). Although the resected brain tissues also include a small portion of the hippocampus, detailed parcellation of hippocampal labels using VH's preoperative MRI in FreeSurfer (Fischl 2012) indicates that the lesion impacts only a small portion of the left hippocampus, preserving 77.5% voxels of the whole left hippocampus while affecting mostly its tail. We identified no lesions in the right hemisphere.

To quantify the function of the resected brain tissue, we retrospectively analyzed VH's iEEG data recorded while he completed a visual categorization task during a seizure-free recording session (Fig. 2A). In this task, participants try to categorize the information content of a presented image into one of four options displayed on the screen on each trial (i.e., ANIMAL, OBJECT, PERSON, PLACE; see further details in Xie et al. 2024). Behaviorally, the participant demonstrated comparable accuracy and response times across stimulus categories (average accuracy = 98.3%; median reaction time averaged across categories = 848 msec).

Figure 2.

Sensitivity of resected brain tissues to scene images in Patient VH. (A) During epilepsy monitoring, Patient VH performed a visual categorization task, identifying images as ANIMAL, OBJECT, PERSON, or PLACE with high accuracy. (B) The electrode probe that later covered resected tissues (pink) had 16 effective contacts, with seven in the parahippocampus (yellow), four in white matter (gray), and six in occipital-temporal cortex (green). (C) Highest classification accuracy (hit—false alarm) occurred for scene images during image presentation until response (median reaction time: 848 msec). Classification for other image categories over time is much attenuated. (D) Classification accuracy for scene images significantly exceeded chance within a task-related time window (bootstrap P < 0.001), whereas other categories did not. No significant category-sensitive patterns were found in lateral recording sites (E and F). (CI) Confidence interval. Each dot in D and F represents decoding results based on shuffled data from a single iteration.

We focused our analyses of the iEEG recordings from electrodes placed in the medial or lateral areas of the resected tissues and avoided those directly in contact with white matter tissues (Fig. 2B). Using trials in which VH successfully categorized the stimulus, we aggregated iEEG power across frequencies and electrodes to decode whether the recorded signals contain any information about the presented stimulus (see Materials and Methods for details). We trained and tested the data separately for each task-relevant time window and for each stimulus category, and quantified the classification accuracy as the difference between classification hit rate and false alarm rate (i.e., hit—false alarm). This metric captures both the sensitivity and specificity in decoding performance, rendering the chance-level decoding accuracy at 0. We find that signals from the medial electrodes in contact with parahippocampal tissue show the highest classification accuracy for scene images toward the end of image presentation until the response (e.g., ∼ 250–750 msec; Fig. 2C). We then examined classification accuracy over a broader time window (stimulus onset to median response time of 848 msec), using all iEEG features within this time window. In the medial contacts, overall classification accuracy again is significantly higher than chance for scene images (classification hit—false alarm for conditions with true vs. shuffled trial labels: 0.12 vs. 0 ± 0.03, bootstrap P < 0.001; Fig. 2D). In contrast, we did not find significant overall classification accuracy for the other stimulus categories. Notably, the aggregated classification accuracy for scene images was about three times that for the second-best category (scene vs. face: 12% vs. 4%). Classification accuracy for scenes is specific to the medial electrode contacts, as we did not find significant decoding for any category in the lateral regions (Fig. 2E,F).

During VH's 36-month visit, we further used fMRI to map selectivity to scenes and faces across his visual cortex following resection (Fig. 1C). We measured BOLD activation in response to viewing blocks of color pictures of faces and scenes (see Materials and Methods). In VH's intact right hemisphere, we observed localized regions of face and scene-selectivity consistent with the conventional FFA and PPA (Kanwisher et al. 1997; Epstein and Kanwisher 1998; Rosenke et al. 2021), respectively. In contrast, in VH's left hemisphere, the posterior portion of PPA appeared intact but the anterior portion was absent due to the lesion. Together, the confluence of evidence from various recording modalities indicates that the excised neocortical regions in VH are genuinely implicated in encoding information related to specific visual categories, particularly scene images, mirroring the functional patterns frequently observed in the PPA (Epstein and Kanwisher 1998).

PPA lesion in VH compromises delayed recognition of scene images

We then examined the impact of the parahippocampal lesion on VH's ability to form and retain novel visual memories. During a 12-month postoperative assessment, VH completed a recognition memory evaluation encompassing diverse visual categories (i.e., SCENE, OBJECT, FACE) (Fig. 3A). In each category, VH studied 80 distinct exemplar images, responding to stimulus-specific queries to ensure engagement (i.e., natural vs. man-made for scenes, edible vs. nonedible for objects, and male vs. female for faces). We tested his memory for half of the images immediately after study, and tested the remainder ∼18 h later (Fig. 3B). We tested categorical conditions in blocks in a random sequence during immediate testing, and mirrored this order during delayed testing. We replicated this protocol with eight neurological patients who did not have lesions in this location (see Supplemental Table S1) in a within-participant manner using identical study and test orders to eliminate any differences that could arise due to testing order. We also replicated this protocol with 145 online participants, each randomly assigned to test in only one of the categorical conditions (i.e., n = 54 for SCENE, n = 43 for OBJECT, and n = 48 for FACE).

Figure 3.

Patient VH shows greater forgetting of scene memory after an overnight delay. (A) During the 12-month follow-up, VH completed a recognition memory test using images of scenes, objects, and faces. For each category, he studied 80 unique exemplar images, during which he answered a stimulus-related question to ensure his engagement (e.g., natural vs. manmade for scenes). Afterward, half of the study images were tested immediately (green) and another half were tested after an 18 h delay (orange). Similar tests were conducted on neurological patients and online participants. (B) Each participant's performance (i.e., hit and false alarm rates, un-circled dots) is shown with averages (circled dots) and standard errors (i.e., horizontal and vertical error bars). The area between the hit rate and false alarm rate contours of the immediate and delayed tests captures the amount of forgetting between these two testing time points. (C) VH's scene memory declined significantly, while object and face memory remained stable. Other participants did not show this pattern. Each line or data point represents the result from one participant. The face photographs in this figure (also in Fig. 4) are sourced from the public domain under Creative Commons noncommercial licenses.

Despite the resection in PPA-affected areas, VH maintains strong immediate scene recognition, even better than his immediate recognition for objects and faces. However, after an ∼18 h delay, VH's scene memory noticeably declines, evident by a larger difference in the area under the curve of the receiver operating characteristics profile (Fig. 3B). To obtain a bias-corrected estimate of this difference, we evaluate VH's visual recognition memory performance as d′ (z[hit] − z[false alarm]) based on the classic signal detection theory (Wickens 2001; Xie and Zhang 2017). VH's forgetting in visual recognition memory may therefore be quantified as the difference in d′ between immediate to delayed tests (Δd′[immediate − delayed]). We find that VH's overnight drop in scene recognition performance (Δd′ = 1.10) is almost twice as large as the drop in performance for objects (Δd′ = 0.54) and for faces (Δd′ = 0.55) (Fig. 3C).

We compared VH's performance to that of the other neurological patients to evaluate the rarity of VH's scene-selective forgetting pattern, which entails directional (one-tailed) tests with a predefined contrast of +2, −1, and −1 for visual memory forgetting for scenes, objects, and faces, respectively (Rosenthal et al. 2000; Corballis 2009; see Materials and Methods for details). We first noted that VH's immediate recognition hit and false alarm rates do not significantly deviate from the centroids of those in other neurological patients based on a test of multivariate outlier detection even with a liberal alpha level of 0.1 (Wilks 1963). However, in the delayed recognition test, in contrast to VH, the other patients exhibit similar levels of forgetting across all three categories [scenes: Δd′ = 0.33 ± 0.08; objects: Δd′ = 0.47 ± 0.09; faces: Δd′ = 0.37 ± 0.10; repeated-measured ANOVA: F(2, 14) = 0.94, P = 0.41]. As a result, the extent to which VH's visual memory forgetting follows the predefined scene-selective contrast is significantly greater than that in other neurological patients in the predicted direction [t(7) = 2.82, one-tailed P = 0.013].

We similarly compared VH's recognition performance to a diverse group of participants recruited through Amazon Mechanical Turk (n = 145). These online participants performed a two-part visual recognition task (∼18 h apart) in one randomly chosen categorical condition. Consistent with the data from the neurological patients, these online participants show no significant difference in overnight forgetting (Δd′[immediate − delayed]) across stimulus categories [one-way ANOVA: F(2, 144) = 0.79, P = 0.46]. Even with the larger sample size, VH's immediate recognition hit and false alarm rates still do not significantly deviate from the centroids of those in the online participants (Wilks 1963). This suggests that VH's immediate recognition memory performance falls within the typical range of this task. Critically, the extent to which VH's visual memory forgetting follows the predefined scene-selective contrast was also significantly greater than that across groups of online participants in the predicted direction (Z = 1.92, one-tailed P = 0.027; see Materials and Methods for this between-subject comparison).

Time-dependent changes in scene memory following PPA lesion

The data so far suggest that VH's selective scene memory deficit becomes evident during the transition from immediate to delayed recognition performance. This inference, however, is based on only two time points for testing visual memory retention. Therefore, to examine how VH's long-term retention of scene images changes over time, at his 36-month follow-up visit, we tested him and a separate group of neurological patients with intact bilateral parahippocampi (n = 12) using a continuous recognition memory task with novel scene images over the course of ∼24 h (Fig. 4A). As a within-subject control condition, we also tested these participants using a continuous recognition memory task with novel face images (Fig. 4B). In this task, we present each participant with scene or face images across different experimental sessions organized into blocks of 100 images from the same visual category (see Supplemental Fig. S1). For each presented image, participants try to decide within four seconds whether they have previously seen the presented image during the task. We calculated participants’ performance as d′, operationalized as the normalized hit rate within a given repetition time-lag bin minus the normalized false alarm rate within the experimental block when participants rendered a memory judgment to account for time-related variation in response bias (Allen et al. 2022). As not all participants completed the task with identical time delays, to improve our measures, we retained the data when participants completed sufficient trials (>10 trials) for estimated recognition memory task performance at a given time bin. Using the metric, we replicated the canonical forgetting curves demonstrated in previous research (Jenkins and Dallenbach 1924; Allen et al. 2022; see Supplemental Fig. S2).

Figure 4.

Patient VH shows rapid forgetting of scene memory compared with other neurological patients. (A) VH (36-month follow-up) and other neurological patients completed a continuous recognition scene memory task over ∼24 h. VH's scene memory significantly declined compared to others after ∼1 h, and the difference persisted for up to ∼24 h. (B) Both VH and control patients also completed a face memory task over ∼24 h, showing similar memory profiles. Dots represent individual data. Error bars represent standard errors of the mean. The face photographs in this figure are sourced from the public domain under Creative Commons noncommercial licenses.

VH demonstrates scene recognition memory performance comparable to that of other neurological patients at short repetition lags (corrected t-test, Crawford and Howell 1998: t(11) = 0.86, one-tailed P = 0.20; Fig. 4A). Similarly, his immediate recognition performance for faces is also comparable to that of the other patients [corrected t-test, Crawford and Howell 1998: t(11) = 0.34, one-tailed P = 0.37; Fig. 4B]. However, VH's scene recognition memory performance rapidly decreases, showing a 95% drop in performance from immediate recognition within 30 sec to the overnight delay on the forgetting curve. This reduction in recognition memory task performance over time is significantly greater than that observed in the other 10 neurological patients who completed the continuous recognition task on Day 2, showing only a 79.7% ± 1.5% overnight drop in performance on the forgetting curve [corrected t-test 35: t(9) = 3.05, one-tailed P = 0.0068]. A complementary resampling analysis confirmed that VH demonstrated a significantly greater reduction in scene memory retention compared to other patients (<30 sec vs. overnight Δd′ = 3.32 vs. 2.26 for VH and others, respectively; bootstrapped one-tailed P < 0.001; Fig. 4A). In contrast, both VH and the control patients show similar forgetting for face images over the course of ∼24 h (<30 sec vs. overnight Δd′ = 2.61 vs. 2.18 for VH and others, respectively; bootstrapped one-tailed P = 0.10; Fig. 4B). Furthermore, the reduction in VH's scene memory overtime relative to other patients was also significantly greater than that for faces (bootstrapped one-tailed P = 0.008). These findings replicate those observed in the 12-month follow-up. The consistency of these results across a 24-month gap highlights the selective and long-lasting impact of parahippocampal lesion on visual memory consolidation for scene images, resistant to potential plasticity or compensation mechanisms (e.g., Celeghin et al. 2017).

Discussion

Enlightened by a case observation, our findings suggest that consolidating scene memories may involve the reactivation of scene-selective neocortical structures across extended temporal intervals. While alternative compensatory mechanisms unexplored in the present study may facilitate initial perceptual encoding, our data highlight that focal cortical lesions can compromise long-term memory consolidation for specific visual details. This scene-selective deficit in visual memory consolidation contrasts with the overall reduction in memory functions often seen in patients with hippocampal lesions (Milner 1968; Nadel and Moscovitch 1997; Squire et al. 2015). Although VH also had a small portion of his left hippocampus resected (i.e., 22.5% voxels of the whole left hippocampus with most of them at the tail), his performance on the recognition tasks demonstrates disproportional deficits for scenes. These data suggest that the impact of VH's small hippocampal resection appears not to exceed the much more pronounced impact of resecting his parahippocampus on scene recognition memory. Furthermore, the increasing decline in VH's scene recognition performance over time, compared to his object or face recognition, suggests that a general memory retrieval deficit for scenes alone does not fully explain these findings. Thus, complementing prior neuroimaging-based observations (Tambini et al. 2010; Shanahan et al. 2018; Collins and Dickerson 2019), our findings argue that category-selective neocortical structures are necessary for long-term memory consolidation of category-specific visual information.

Our case study here, therefore, provides several unique insights into the role of the neocortex in memory consolidation, despite various well-recognized challenges often associated with case reports (Vaidya et al. 2019). First, identifying specific memory impairments associated with neocortical lesions has been challenging, as the precise information encoded by the excised tissue often remains unknown in many brain lesion cases (De Renzi 1982; Vaidya et al. 2019; Xie et al. 2023a). The standard approach to this challenge has been to compare the lesion area with a probabilistic functional atlas. However, chronic seizure activity and its spread can disrupt the function of these typical category-selective neocortical areas (Diamond et al. 2024), making it difficult to directly associate these structures with their functions. Here, our data tackle this uncertainty by analyzing iEEG data recorded from the subsequently resected brain tissues during a visual categorization task. Furthermore, we verified these results in VH's post-lesion brain using fMRI—a complement to the iEEG results at a different spatiotemporal scale. These analyses found scene-selective information coding in the medial portion of the resected brain tissues, consistent with the observed deficits in scene-selective memory consolidation in this case.

Second, because pre-surgery behavioral data can be hard to obtain due to scheduling challenges in a clinical setting, many studies have to rely on the comparison of task performance between individuals with brain lesions and control participants (Vaidya et al. 2019). This presents some interpretational difficulties, as neurological cases often involve various complications (e.g., seizures, medications, etc.). To mitigate this, we have compared VH's visual memory across different categories within himself, in addition to contrasting his data with other neurological patients who had similar clinical circumstances within the same clinical setting and with online participants from the general population. Our data consistently point toward a distinctive deficit in scene memory consolidation for VH. This case study, therefore, suggests a previously less characterized time-dependent memory deficit following PPA damage. Future research with a larger control cohort using lesion-symptom mapping techniques may further evaluate these memory functions.

Hence, despite various known inferential challenges based on a single case, the findings from this current case, like many prior case reports (Jacobs et al. 2012; Vaidya et al. 2019), pose intriguing questions for future research. First, it will be important to articulate the precise mechanisms through which category-selective neocortex is involved in memory consolidation. Previous research suggests that after initial encoding, enhanced memory reinstatement is linked to the coordination between rapid local oscillatory signals (such as 80–120 Hz ripples) in the neocortex and those in the medial temporal lobe (Norman et al. 2019; Vaz et al. 2019). It is plausible that the reactivation of category-selective neocortex is facilitated by and directly driven by hippocampal ripples. However, thus far, evidence supporting this conjecture remains scarce, and studies examining such coordination typically only focus on memory retention over a relatively short delay (Norman et al. 2019; Vaz et al. 2019). Examining the relation between hippocampal and neocortical activity, especially during overnight sleep, will be a crucial test for the proposed neurophysiological mechanism concerning content-specific memory modulation in the human brain.

Second, it is also important to identify the specific region within the category-selective neocortex that is pertinent to memory consolidation. Earlier studies have indicated that even within conventional category-selective regions of the neocortex, such as the PPA, there could be a potential functional division between perception and memory (Aminoff et al. 2013), potentially divided between the posterior and anterior parts of the PPA, respectively (Steel et al. 2021, 2023). While our current data do not allow for a direct investigation of this possible distinction, the fMRI results in VH's post-lesion brain generally align with this distinction. For instance, while the posterior part of the left PPA remains intact in VH, the anterior part is missing due to the lesion. This stark contrast with the typical scene-selectivity on the intact hemisphere raises an intriguing question regarding how the anterior, as opposed to the posterior part of the PPA, contributes to scene-selective visual memory consolidation compared to scene perception. Future research involving direct recordings from both the anterior and posterior PPA in intact brains using a long-term memory consolidation paradigm may help address this question.

Lastly, it is important to identify the functional plasticity and compensation mechanisms following unilateral PPA resection. While VH has undergone a well-characterized resection of the PPA, his short-term memory for scene images remains intact and as good or better than his memory for other stimulus categories. This is evident in VH's immediate recognition performance at the 12-month visit, as shown in Figure 3, and his performance in the continuous recognition task during short delays at the 36-month visit, as depicted in Figure 4. VH's scene memory deficits only become apparent after a few hours, particularly overnight at both the 12- and 36-month visits, suggesting a potential difficulty in transferring these temporary scene memories into long-term storage. However, it remains unknown how VH forms the temporary scene memories in the first place. Various hypotheses can be proposed regarding compensatory mechanisms involving the intact hemisphere or other cortical areas within the same hemisphere (Kolb and Whishaw 1989; Holtmaat and Svoboda 2009). Future research should further investigate these mechanisms to better understand the plasticity and resilience of the human brain, for example, by evaluating cases with bilateral lesions (Epstein et al. 2001).

In summary, our data demonstrate that damage to scene-selective neocortical structures can compromise the consolidation of visual memories related to the corresponding category. As such, the reactivation of category-selective neocortical areas may be required to convert transient visual perceptual experiences into lasting visual memories.

Materials and Methods

Participants

Patient VH (right-handed, male, 36 years old at initial assessment in December 2018, IQ = 103) participated in the study at the Clinical Center of the National Institutes of Health (NIH, Bethesda, Maryland, USA) for the treatment of epilepsy. He underwent a surgical procedure, during which platinum recording contacts were implanted in-depth via stereo-electrode probes to localize epileptogenic brain regions (see Case history for details). The location of these recording sites was clinically determined by the attending neurologists and neurosurgeons. During a recording period when no seizure events were detected, VH completed a visual categorization task from which we identified category-selective coding of the resected brain tissue (detailed in Procedure). Following his brain surgery, VH returned in November 2019 and in September 2022 to complete additional behavioral testing and functional neuroimaging.

Furthermore, we recruited 20 neurological patients at the NIH as control participants (10 females; 37.15 ± 2.04 [mean ± s.e.m.] years old; Wechsler IQ: 85.06 ± 3.82; see Supplemental Table S1). Inclusion criteria of these other neurological patients include the (1) consent and (2) availability to participate in at least one of the same tasks completed by VH, as well as (3) no prior lesion in bilateral parahippocampal cortices. As an additional control group for the immediate and delayed visual recognition memory tasks, 145 online participants were recruited through the Amazon Mechanical Turk (mTurk) experimental platform (35.64 ± 0.78 years; 65 males; $10/h for monetary compensation). Inclusion criteria of these online participants include the (1) consent and (2) completion of both the immediate and overnight delayed memory tests, along with (3) correctly responding to various check questions to ensure participants understood and followed task instructions (Xie et al. 2020b). The timeline of these study events and participants’ involvement in different parts of this study is summarized in Figure 1A. The Institutional Review Board (IRB) at the NIH approved the research protocol, and informed consent was obtained from the participants or their guardians.

Case history

VH's onset of seizures was at age 19. He played football and other contact sports for years and had experienced a few concussions. There were no other identifiable causes or precipitating events for the onset of his seizures. In 2018, VH underwent surgical treatment at the NIH and received left temporal corticectomy and posterior hippocampectomy. During the preoperative interview, VH reported word-finding difficulties and decreased memory abilities. These complaints were confirmed by standard neuropsychological evaluation administered in 2017, during which no motor abnormalities, speech difficulties, or emotional distress were noted. Test results suggested that VH's general intellectual function fell within the average range, with lower-than-average verbal abilities and average visuospatial abilities and perceptual reasoning. VH showed normal color vision, and his digital span was within normal limits. He obtained an associated degree, reported no history of learning difficulties, and was employed part-time.

Procedure

Visual categorization task

To determine the relevance of the recorded brain tissue in representing higher-level visual information, VH completed a visual categorization task during iEEG monitoring implemented using PyEPL (Geller et al. 2007; Xie et al. 2024). Each trial began with a single image at the center of a 15-in laptop screen for 500 msec. VH was instructed to press one of the four arrow keys corresponding to the text label that best classified the present image (i.e., ANIMAL, OBJECT, PERSON, PLACE) (see Fig. 2A). The task would not progress until the participant made a choice, thereby favoring accuracy over response time. Following the response, a blank screen appeared for 200 msec before the presentation of the next image. Additionally, VH passively viewed a noise image with scrambled pixels at the beginning of each experimental block for 1000 msec as baseline activity that could be used for subsequent normalization. Feedback regarding response accuracy was provided at the end of each block (e.g., “56 out of 60 correct”).

For each taxonomic category (ANIMAL, OBJECT, PERSON, PLACE), we selected 60 unique images and randomly chose 15 of them from each category within an experimental block. As a result, each experimental block consisted of a total of 60 images across the four taxonomic categories. To control for individual features that could aid in visual categorization, all of these images were resized, cropped, and lightly phase shifted. We used the SHINE toolbox to balance the set of images for their luminance, contrast, and spatial frequency (Willenbockel et al. 2010). Ultimately, each presented image occupied half the size of a 15 inch laptop computer screen and appeared in grayscale (see Fig. 2A, for examples). One to 3 days before surgical implantation, VH was given a practice round with a reduced trial count. During the recording, VH completed 24 experimental blocks, yielding a total of 1440 trials with 360 trials per category. Task performance is quantified as the categorization accuracy and reaction times.

Immediate and delayed visual recognition memory task

During VH's 12-month follow-up visit, he was invited to complete a visual recognition memory task over two experimental sessions on consecutive days, implemented using Psychtoolbox (Brainard 1997). We subsequently repeated the same procedure with another eight neurological patients who had intact bilateral parahippocampi and with 145 online participants who completed this task with minor modifications via Amazon Mechanical Turk (see below).

On Day 1, participants began by studying a list of 80 unique images depicting scenes, faces, or everyday objects (see Fig. 3A). Each image was presented at the center of a laptop/computer screen for 3000 msec, followed by an interstimulus interval of 1000 msec. The image, presented in color and in a square format, occupied 35% of the computer screen. To ensure active encoding of the presented images, participants were instructed to classify the images according to a property of the stimulus category—whether the scene was indoor or outdoor, whether the face was male or female, and whether the object was edible or inedible. Participants were asked to make a two-alternative force-choice response within 4000 msec using the left and right arrow keys. After displaying the 80 study images, 40 of them were randomly chosen from the list for an immediate recognition memory test. Another 40 novel images from the same category were added and intermixed with the original 40 study images in a random order. Participants were asked to make a new versus old recognition memory judgment for each test image, using the left and right arrow keys, respectively. The image would remain on the screen until the participant's response, followed by a 1000 msec intertrial interval, emphasizing accuracy over speed. VH completed three blocks of immediate recognition memory tests, separately for faces, scenes, and objects in the mentioned order with short breaks in between each immediate study-and-test block. To rule out study and test order as a potential confound, the eight neurological patients completed the immediate recognition memory tests using the same study and test orders for images within each block and for categorical conditions across blocks as VH. The online participants were randomly assigned to one of the three categorical conditions, using the same within-block image study and test order as VH.

On Day 2, which occurred ∼18–24 h after Day 1's testing (e.g., if Day 1 testing was conducted from 3 p.m., Day 2 testing took place between 10 a.m. and 3 p.m.), participants completed three blocks of the delayed recognition memory task, in which each block contained the remaining 40 study images from one of the stimulus categories that were randomly intermixed with another 40 novel images from the same category. Participants once again made untimed new versus old recognition memory judgments for each presented test image using the left and right arrow keys, respectively. VH completed this delayed recognition test across stimulus categories in the same testing order as his immediate recognition memory test. Other neurological patients also followed this testing order both within each block and across stimulus categories. Online participants tested images from one of the stimulus categories using the same within-block image test order as VH. Participants’ performance is quantified as the hit and false alarm rates in this recognition memory task. As the number of new and old images were balanced in this task, we also calculated the recognition memory discriminability (d′ = Z[hit] − Z[false alarm]) separately for immediate and delayed memory tests based on the classic signal detection theory model (Wickens 2001).

Continuous visual recognition memory task

During VH's 36-month follow-up visit, he was invited to complete a continuous visual recognition memory task across experimental sessions on consecutive days. We subsequently repeated similar procedures with reduced trial counts in another 12 neurological patients who had intact bilateral parahippocampi. On each trial, participants saw an image at the center of a 15 inch laptop screen, which occupied 35% of the presentation screen in a squared form. Each image remained on the screen for 2000 msec, followed by a 2000 msec interstimulus interval. For each presented image, participants reported whether they have seen this image before or not during the experiment using the left and right arrow keys, respectively, within a maximal of 3000 msec response time window. Following the response, participants heard a “cha-ching” sound (1000 msec) upon correct recognition. No sound played if participants made a wrong judgment. Participants completed this task in blocks of 100 images from either the scene or face category, taking ∼7 min per block. Within each block, 10% of images repeated with a short lag of one to 10 images in between and 20% images repeated with longer lags that were uniformly distributed from possible image lags (Supplemental Fig. S1A). This manipulation ensures that the likelihood of repetition is 30% across blocks (Supplemental Fig. S1B), mitigating the potential shift in response criteria if repetitions were more likely toward the end of the experiment (Allen et al. 2022). Within each experimental session, participants completed one to three blocks as time permitted, with a resting period between blocks taken at their own pace. We spread out different experimental blocks and sessions on consecutive days without specific time constraints. This naturally creates temporal jittering for image repetition at different time lags, from within seconds to up to ∼24 h (see an example in Supplemental Fig. S2A). During his visit, VH first completed 10 blocks of the task using scene images on consecutive days. Subsequently, he completed another 10 blocks of the task using face images on separate consecutive days. This yielded a total of 1000 trials for each stimulus category. Similar procedures were implemented for other neurological patients, except that they completed about half of the experimental trials across ∼24 h relative to those completed by VH (483 ± 24 trials per category, ranging from 400 to 600 trials per category) with intermixed scene and face blocks in each experimental session.

To quantify participants’ recognition success with images repeated at different time-lag bins (i.e., <30 sec, 0.5–7 min [within a block], 7–30 min, 0.5–2, 2–6, 6–30 h [overnight]), we identified trials repeated at each time-lag bin in each experimental block and calculated the hit rate over these trials. To correct for the tendency of false recognition (i.e., reporting a novel image as “seen before”) that may vary across experimental blocks (Allen et al. 2022), we computed the block-specific false allarm rate (Supplemental Fig. S2B). To ensure a reasonable estimate of the time-lag-specific hit rate, we retained data from a time-lag bin only when it contained >10 trials. Among the 12 other neurological patients, two did not complete enough trials to evaluate overnight recognition memory performance. Participants’ overall corrected recognition performance is estimated as d′, which is operationalized as the average normalized difference between time-lag-specific hit rate across experimental blocks and block-level false alarm rates (Supplemental Fig. S2C). This metric reveals time-lag-specific memory strength while accounting for potential response bias during memory judgments collected at each experimental block. This approach replicates the classic findings of the forgetting curve (Jenkins and Dallenbach 1924; Allen et al. 2022).

Pre-lesion iEEG recordings and analysis

We recorded iEEG signals sampled at 1000 Hz using a data acquisition system from Blackrock Microsystems. The depth electrodes featured eight to 16 contacts spaced 2–5 mm apart, with a diameter of 0.8 mm. We localized the electrode placement by co-registering postoperative CT scans with preoperative MRI T1 images, using a previously established method (Trotta et al. 2017). We aligned the postoperative MRI T1 image that contains the lesion with the preoperative MRI T1 image to identify the scope and location of the lesion area. From these aligned images, we identified 16 electrode contacts on a single electrode probe within the lesion area (Fig. 2B). These contacts were distributed as follows: #1–7 covered parahippocampal tissues on the medial side, #8–11 were in contact with white matter tissues, and #12–16 were positioned on the lateral side, covering occipital-temporal cortical tissues. To compare the lesion's location with the probabilistic functional atlas of scene/place and face-selective cortices (Rosenke et al. 2021), we generated a lesion mask aligned with VH's preoperative MRI T1 image and normalized it to the MNI brain (Fig. 1B). Furthermore, we validated the lesion area using subject-specific parcellation from FreeSurfer based on the preoperative MRI T1 image (Fischl 2012).

Our primary interest was the information coding of the identified electrode contacts that covered neocortical tissues in the lesion area. We therefore proceeded with our iEEG analysis using the medial (#1–7) and lateral (#12–16) recording sites, following the preprocessing and analysis steps outlined below. First, we attempted to identify and reject noisy contacts with an average amplitude or variance >3 standard deviations of the mean estimates from all available electrode contacts in the recording session. None of the identified electrode contacts were rejected based on this criterion. Second, we removed slow fluctuations using local detrending and eliminated line noise at 60 and 120 Hz. Subsequently, we bipolarly referenced the remaining iEEG traces using every two adjacent contacts on the same electrode probe. These bipolar-referenced signals are henceforth referred to as electrodes. Third, we discarded additional electrodes and trials that exhibited excessive kurtosis or variance during individual trial epochs, as previously detailed (Wittig et al. 2018; Xie et al. 2020a, 2023a). These combined steps were undertaken to eliminate iEEG artifacts potentially influenced by interictal activity, movement-induced artifacts, or transient electrical perturbations from external sources (Wittig et al. 2018; Chapeton et al. 2019; Xie et al. 2020a). In the end, we retained 97.8% of the trials. Moving forward, we quantified the spectral power of these cleaned iEEG signals in epochs from −2000 to 4000 msec following stimulus onset through wavelet analysis (across 30 logarithmically spaced values ranging from 3 to 150 Hz, with a wave number of 6). The 1000 msec data buffers at the start and end of each epoch were excluded. For every frequency and electrode, we extracted z-scored power values based on the mean and standard deviation of power during the presentation of a noise image. We smoothed the data using a 200 msec sliding window, with steps of 20 msec. Finally, we averaged the power values within each time window across frequencies, organized into conventional frequency bins (θ: 4–8 Hz, α: 8–16 Hz, β: 16–32 Hz, γ: 32–70 Hz, high-frequency broadband: 70–150 Hz). These iEEG features were then vectorized across electrodes and frequency bins for each time window, creating a feature vector for subsequent analysis (Yaffe et al. 2014; Xie et al. 2023b).

To determine whether the iEEG features extracted from the recorded signals in the later lesion areas contain information about the presented image categories, we constructed linear logistic regression classifiers with regularization (Du et al. 2017; Xie et al. 2024). We utilized both the vectorized iEEG features for each time window (feature length = # electrodes × 5) and the aggregated iEEG features across all task-relevant time windows (from stimulus onset to the median reaction time of 848 msec; feature length = # electrodes × 5 × # time windows). For each image category, we employed trials in which participants accurately identified the category. We implemented a 10-fold cross-validation procedure to estimate the classification accuracy using a one-versus-all model. Each hold-out set comprised ∼10% of all trials, uniformly distributed throughout the experimental session (see Supplemental Fig. S3). Within each category, the classifier's performance was quantified as the percentage of hold-out trials with correct predictions minus the percentage of hold-out trials with incorrect predictions for that category (i.e., hit—false alarm). To mitigate issues related to multiple comparisons, we conducted statistical tests solely on classification outcomes based on the iEEG features aggregated across all task-relevant time windows. To accomplish this, we established an empirical chance-level classification accuracy by training and testing the classifier on data with shuffled trial labels.

Post-lesion neuroimaging and analysis

fMRI data acquisition

During VH's 36-month follow-up visit, we used fMRI to localize selectivity to scenes and faces in VH's visual cortex (see Fig. 1B). MRI data were acquired with a 3.0T GE Discovery MR750 MRI scanner and 32-channel head coil at the Functional Magnetic Resonance Imaging Core Facility (FMRIF) at the NIH in Bethesda, MD. A high-resolution T1-weighted structural MRI scan (3D-MPRAGE sequence, 1 × 1 × 1 mm3 voxel size, in-plane matrix size: 256 × 256, 172 slices, TR = 7.66 msec, TE = 3.48 msec, FA = 7°) was collected. Functional scans were collected with a three-echo echo-planar sequence (TR = 2000 msec, TEs = 12.5, 27.7, and 42.9 msec, flip angle = 75°, 64 × 64 acquisition matrix, in-plane resolution = 3.2 × 3.2 mm, slice thickness = 3.5 mm, 30 slices). Visual stimuli (size: 5° × 5° of visual angle) were displayed using a flat-panel MRI-compatible 32″ Cambridge Research Systems BOLDscreen with a resolution of 1920 × 1080 and a viewing distance of 2.5 m. Behavioral responses were collected using an MRI-compatible button box.

Functional localization of scenes and faces

Two functional localizer runs (5.3 min duration each) were used to define scene and face-responsive regions. The functional localizer stimuli were color photographs of 162 faces and 162 scenes (600 × 600 pixels). Each localizer run started and ended with a 16 sec fixation period, during which a black fixation cross was centered on a gray screen. The localizer used a block design, alternating between 16 sec blocks of faces and scenes. There were nine blocks of each stimulus type, for 18 blocks total. Within each 16 sec block, 20 stimuli were shown one at a time on a gray background in random order for 300 msec with a 500 msec interstimulus interval. To maintain the participants’ attention throughout the localizer runs, we used a standard 1-back task. VH was instructed to press a key when they saw the same image presented twice in a row. In each block, 18 unique stimuli were shown, and two stimuli were randomly selected to be repeated for the 1-back task. There were two 1-back trials per block, and they were spaced so that one repeated image occurred at a random interval in the first half of the block, and the second repeated image occurred at a random interval in the second half of the block. Each of the 162 unique stimuli in each stimulus class (faces and scenes) was only shown once per run, except for the images repeated for the 1-back trials. The images were randomly allocated to blocks each time the code was run. Mean task accuracy across both runs was 72.22% ± 8.34%.

fMRI analysis

fMRI data were preprocessed using the AFNI software package (Cox 1996). Functional images were slice-time corrected, motion-corrected, and co-registered to VH's in-session anatomical volume using the customizable script afni_proc.py. Spatial smoothing of 4 mm full-width at half-maximum was applied to the functional localizer runs. The multi-echo EPI data was combined into a single time series using AFNI's optimally combined algorithm. All analyses were conducted in VH's native brain space. For data visualization (see Fig. 1B), the results of the generalized linear model contrast (faces—scenes) were overlaid as t-maps on the anatomical volume.

Statistical analysis

We aim to assess the rarity or abnormality of VH's visual memory function by comparing his task performance relative to a normative or control sample, which typically involves a one-tailed test (Crawford and Howell 1998; Crawford and Garthwaite 2012). In the current case, we are motivated to examine scene-selective deficits in VH, which entails a composite score across stimuli categories. Yet, a direct comparison of absolute task performance at a given time point could be influenced by individual- or stimulus-level differences, making across-category comparisons less effective. To address this, we calculated a forgetting score by subtracting participants’ delayed memory task performance from their immediate memory task performance, thereby isolating the amount of forgetting over time for each stimulus category. This baseline-corrected score provides a more accurate measure of the differences in forgetting across stimulus categories. For instance, to compare VH's 12-month visual recognition task performance relative to other neurological patients, a scene-selective forgetting profile would predict participants’ overnight forgetting for scenes, objects, and faces should follow predefined contrast weights (λ) of +2, −1, and −1, respectively. A within-subject composite score (Rosenthal et al. 2000), Formula, may therefore be estimated for subsequent comparisons of task performance (x) between VH and other patients using the previously established methods (Crawford and Howell 1998; Crawford and Garthwaite 2012), Formula where Formula and SL are the mean and standard deviation of the composite L scores across control samples. As the current estimate of participants’ task performance at each time point is highly reliable within individuals (see split-half analysis in Supplemental Fig. S4 based on the procedure previously described in Xie et al. 2020c), this contrast analysis can be evaluated using a one-tailed P-value to determine whether the predefined contrast weights significantly predict the visual memory forgetting profile across categories in a single-case relative to others under the same task conditions (Crawford and Howell 1998; Rosenthal et al. 2000; Crawford and Garthwaite 2012).

In the context where VH's forgetting for scenes/objects/faces is compared with the average performance across different groups of online participants, we extend this contrast analysis using conventional standardized difference scores between VH and online participants in each corresponding condition. We assume that the online participants reflect the population of interest in the current comparison—resembling a fixed-effect model (Rosenthal and Rosnow 2008). As such, for each category i, we estimate how extreme VH's forgetting is relative to the online participants using the mean, μi, and standard deviation, si, of the control sample: Formula

The extent to which the scene-selective forgetting contrast in VH (λ: +2, −1, −1 for scene, object, and faces, respectively) applies to other online participants can then be evaluated by a contrast analysis of the standardized scores across experimental groups (Corballis 2009), which is approximated as a test of the weighted sum of standardized scores (Rosenthal et al. 2000; Rosenthal and Rosnow 2008; Xie et al. 2023c): Formula

In addition to these contrast tests, we incorporate results obtained from resampling or permutation tests where appropriate (e.g., when estimating classification outcomes against empirical chance performance; Good 2013). All P-values associated with case-control comparison and predefined contrast are evaluated based on one-tailed tests (Crawford and Howell 1998; Rosenthal et al. 2000; Crawford and Garthwaite 2012). If not otherwise specified, P-values are reported as two-tailed.

Data access

Except where otherwise noted, computational analyses were performed using custom written MATLAB (MathWorks) scripts, available upon request. Processed data used in this study are available at: https://research.ninds.nih.gov/zaghloul-lab/downloads.

Competing interest statement

The authors declare no competing interests.

Acknowledgments

We thank all the participants who selflessly volunteered their time for this study. We also thank John Wittig Jr., Julio Chapeton, and Joshua Diamond for insightful comments on the study. This work was supported by the Intramural Research Programs of the National Institute for Neurological Disorders and Stroke (ZIA-NS003144, PI: K.A.Z.) and the National Institute for Mental Health (ZIA-MH002909, PI: C.I.B.). W.X. is supported by the National Institutes of Health (NIH) Pathway to Independent Award (R00NS126492). We are indebted to all participants who have selflessly volunteered their time to this study.

Author contributions: W.X. conceptualized the study and crafted the initial draft with guidance from K.A.Z. and C.I.B. S.G.W. acquired and analyzed the fMRI data, contributing to manuscript preparation. J.L. managed data and contributed to manuscript preparation. S.J. contributed to acquiring the online data. O.F., M.B., A.P., and A.P.T. contributed to data collection for control patients. S.K.I. oversaw iEEG data acquisition and patient care throughout the study. K.A.Z. performed all surgical procedures and supervised the study. K.A.Z., C.I.B., and S.G.W. provided crucial editorial input to the paper.

Footnotes

  • Received July 18, 2024.
  • Accepted March 6, 2025.

This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the first 12 months after the full-issue publication date (see http://learnmem.cshlp.org/site/misc/terms.xhtml). After 12 months, it is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.

References

| Table of Contents