The influence of categorical stimuli on relational memory binding
- 1Department of Psychology, University of Nebraska-Lincoln, Lincoln, Nebraska 68588, USA
- 2Center for Brain, Biology, and Behavior, University of Nebraska-Lincoln, Lincoln, Nebraska 68588, USA
- 3Department of Psychology, Binghamton University, Binghamton, New York 13902, USA
- 4Department of Psychology, University of Illinois at Urbana-Champaign, Champaign, Illinois 61820, USA
- 5Carle Illinois College of Medicine, University of Illinois at Urbana-Champaign, Champaign, Illinois 61820, USA
- Corresponding author: hschwarb2{at}unl.edu
Abstract
Binding of arbitrary information into distinct memory representations that can be used to guide behavior is a hallmark of relational memory. What is and is not bound into a memory representation and how those things influence the organization of that representation remain topics of interest. While some information is intentionally and effortfully bound—often the information that is consistent with task goals or expectations about what information may be required later—other information appears to be bound automatically. The present set of experiments sought to investigate whether spatial memory would be systematically influenced by the presence and absence of distinct categories of stimuli on a spatial reconstruction task. In this task, participants must learn multiple item-location bindings and place each item back in its studied location after a short delay. Across three experiments, participants made significantly more within-category errors (i.e., misassigning one item to the location of a different item from the same category) than between-category errors (i.e., misassigning one item to the location of an item from a different category) when categories were perceptually or semantically distinct. These data reveal that category information contributed to the organization of the memory representation and influenced spatial reconstruction performance. Together, these results suggest that categorical information can influence memory organization, and not always to the benefit of overall task performance.
The fidelity of memory is intrinsically tied to the content of what is being remembered. When it is possible to organize the content into meaningful units, the grouping together of the elements of an experience has long been shown to benefit memory performance. For example, there are classic reports of mnemonists remembering 80+ digit strings by grouping numbers into, for example, running times (Ericsson et al. 1980; Staszewski 1990), thus reducing the memory burden by organizing the information into smaller, meaningful chunks of to-be-remembered information. This chunking effect is evident across a multitude of list learning studies (e.g., Miller 1956; for reviews, see Shiffrin and Nosofsky 1994; Norris and Kalm 2021). And while chunking or organizing information into related clusters has well-established benefits (Bousfield and Cohen 1953), it can also lead to list intrusion errors, often in the form of related, but incorrect items. Such intrusions suggest that list content can also influence a memory representation leading to predictable patterns of memory errors (Deese 1959; Roediger and McDermott 1995). For example, visually similar but incorrect items are often substituted for studied items at retrieval (e.g., Pantelis et al. 2008) and semantic distortions are also common, such that individuals falsely recall/recognize items that are related in meaning to studied items (Romney et al. 1993; Howard and Kahana 2002). Remembering lists of information is certainly one specific mnemonic exercise that is prevalent both in and out of the laboratory. However, more common are events in which there are, simultaneously, multiple pieces of related and unrelated information and both the items and also the relations among those items need to be remembered (Cohen and Eichenbaum 1993; Eichenbaum and Cohen 2001). The specific organizational features of these relational memory representations and the types of information that get bound continue to be explored. In the current work, we investigate the ways in which including distinct categories of stimuli can influence the organization of spatial relational memory representations, which, in turn, can alter memory performance.
For nearly a century, researchers have understood that memory is rarely a faithful reproduction of an experienced event. Errors are typical, and the types of errors made can inform us about how information is organized in memory (Bartlett 1995; Schacter 2012; Schacter et al. 2022). Memories are vulnerable to distortion and the elements of an experience itself can contribute meaningfully to this distorting of memory (Loftus and Pickrell 1995; Loftus 2003). Indeed, similarity among items can facilitate generalization, but also create interference, thus contributing to both improved and impaired memory performance depending on task goals (Nelson et al. 2013). One example of mnemonic enhancement is that memory recall is improved for semantically homogenous lists (e.g., tweed, corduroy, gabardine, denim) compared to semantically heterogenous lists (e.g., veal, sofa, arrow, dollar) (Cohen 1963; Romney et al. 1993). Conversely, an example of mnemonic impairment is that novel stimuli (e.g., a beach scene) are often (incorrectly) remembered as having been previously seen when they are very similar to items that were part of the study set (e.g., a different beach scene), and this is true for both perceptually and semantically similar stimuli (e.g., Carpenter et al. 2023; Pantelis et al. 2008; Sternberg 1966).
Fuzzy trace theory provides an interpretive framework for understanding why similarity among stimuli can sometimes improve and sometimes hinder memory outcomes. According to this theory, in a given episode, verbatim information (i.e., precise, literal information) and gist information (i.e., information about the general meaning) are processed and stored separately (Reyna and Brainerd 1995; Reyna et al. 2016). When an episode includes very similar items and precise information about any particular item is not required, gist memory may be sufficient and having highly similar stimuli will facilitate memory outcomes. However, if specific detailed information is necessary, gist information is not sufficient and without a strong verbatim memory trace, memory will be impaired when competitors are very similar. Associative/relational memory tasks are typically of the latter variety. For example, Greene and Naveh-Benjamin (2020) and Carpenter et al. (2023) both showed participants face-scene pairs. In both studies, scenes came from distinct categories with multiple exemplars (e.g., parks, malls, kitchens, offices), and participants were asked to make an old-new judgment about old and new pairs at test. Participants frequently endorsed a new pair as “old” when the scene was highly similar (i.e., came from the same category) to the scene that had been studied at test. Categorical “gist” information alone was not sufficient to support accurate task performance.
The literature described thus far clearly demonstrates that the content of a memory set can influence the specific items that are remembered and misremembered and how categorical information (either perceptual or semantic) can influence item memory. Stepping outside of laboratory settings, it is obvious that item information does not exist in isolation. Rather, items are part of richer scenes or environments and appear in different locations in space and at different moments in time. That raises questions of what sorts of information are bound and how those bindings are organized in memory (Burgess et al. 2001; Maguire et al. 2016), emphasizing particularly the binding of relational information in scenes and events and whether such binding occurs only intentionally or occurs automatically.
Smith and Milner (1981) presented individuals (with and without hippocampal damage) with 16 toys at different locations on a table and asked the participants to estimate the price of each toy. The toys were then removed, and two surprise memory tests were administered. First, participants were asked to name the items that had been presented (item-memory test). Then the toys were returned to the participant and they were asked to place them on the table where they had originally been presented (spatial-memory task). On the item-memory test, participants without brain injury remembered over 50% of the toys (chance = 6.3%) when tested immediately. On the spatial-memory test, the average displacement error (distance between where a toy had been studied and where it had been placed at test) for participants without brain injury was ∼ 6.4 cm suggesting that the participants had automatically bound the spatial information for each toy despite the location having no relevance to the task: price estimating. These data suggest that seemingly irrelevant dimensions in an episode are still often bound into the memory representation and can be accessed and used later.
In the current set of studies, we investigated the ways in which categorical information (not directly relevant to the task goals, but inherent to the stimuli) may influence the organization of spatial relational memory representation and influence memory outcomes. This work takes advantage of the spatial reconstruction paradigm (Smith and Milner 1981; Horecka et al. 2018), a standard relational memory task in which participants study some number of stimuli in different locations on a computer screen and then, after a brief delay, must recreate that study display by putting each stimulus back in its studied location in an open field without fixed locations. In this paradigm, a variety of errors are typical, including mistakenly putting a stimulus in a location where a different stimulus had been studied (i.e., a misassignment). These misassignment errors reflect a breakdown in the fidelity of relational memory. Here we modified the spatial reconstruction task to include stimuli from different perceptual (i.e., blue vs. yellow) and/or semantic (e.g., birds vs. musical instruments; cats vs. dogs) categories. Of note, the task was to learn information about specific item-location pairings, and while the items happen to be drawn from distinct categories, that category information was not explicitly informative to spatial position on the screen. The question was, then, does categorical information influence the organization of a spatial memory representation; and if so, does such influence alter spatial task performance? Each reconstruction was assessed for accuracy with a particular focus on stimulus misassignments, examining whether a misassignment was more likely, equally likely, or less likely to occur among members of the same category versus members of distinct categories. In this set of experiments, we investigated categories that were defined either by their visual similarity (color; Experiment 1), their semantic similarity (taxonomic; Experiment 2), or both (taxonomic categories with manipulated visual similarity; Experiment 3). Together, these data can inform us about (1) whether the conglomeration of information that gets intentionally and/or automatically bound influences the organizing principal of the remembered representation and (2) if that organization shapes behavior.
Results
Experiment 1
The purpose of Experiment 1 was to investigate the impact of perceptual stimulus categories on the organization of relational memory representations using the spatial reconstruction task. In this experiment, categories were defined based on the color of the stimulus (i.e., yellow or blue). On each trial, participants studied the location of six novel, abstract shapes that were either all blue or all yellow (one-category trials) or half blue and half yellow (two-category trials; Fig. 1). We first evaluated whether the presence of one versus two categories affected the fidelity of memory. Next, we considered the nature of memory distortions specifically on two-category trials by comparing the number of stimuli misassigned to a location that had been occupied by a stimulus from the same perceptual category or the other perceptual category.
Example study and test trial in the spatial reconstruction task. First, six stimuli were presented on the screen. All stimuli were either all from the same category or three stimuli were each from two distinct categories. The stimuli then disappeared for a brief 4 sec delay before reappearing at the top of the computer screen. Participants had unlimited time to use the mouse and drag each stimulus back to its remembered location. Participants could rearrange the stimuli as many times as they wanted to and indicated that they were finished with their reconstruction by pressing the space bar.
First, we assessed whether overall performance was impacted by having one versus two categories. Average relational memory accuracy (i.e., how frequently an item was placed in its studied location) was 56.8% (SD = 10.2%) on one-category trials and 57.3% (SD = 11.9%) on two-category trials. The difference was not significant, t(49) = −0.41, P = 0.685, d = 0.06. The average misplacement error (i.e., how far an item was placed from its studied location in pixels) was 190.6 pixels (SD = 60.4) on one-category trials and 191.6 pixels (SD = 64.0) on two-category trials. The difference was also not significant, t(49) = −0.23, P = 0.821, d = 0.03. The average total number of misassignments was 20.2 (SD = 13.3) across one-category trials and 19.3 (SD = 14.1) across two-category trials (Fig. 2). The difference in total misassignments between one- and two-category trials was not significant, t(49) = 0.97, P = 0.338, d = 0.14. In short, overall spatial memory measures of reconstruction accuracy did not differ between one- and two-category trials.
Experiment 1 misassignment data. Total number of misassignment errors for one-category and two-category trials (blue bars) and misassignment errors for two-category trials separated into within-category misassignment (WCM) errors and between-category misassignment (BCM) errors (green bars). Plotted error bars indicate SEM. (***) P < 0.001.
Next, we looked at the types of misassignments made in two-category trials. The average total number of within-category misassignments (WCM) was 13.5 (SD = 10.2) and the average total number of between-category misassignments (BCM) was 5.8 (SD = 6.1). The difference between WCM and BCM was significant, t(49) = 5.96, P < 0.001, d = 0.65.
The results from Experiment 1 demonstrate that salient color categories can contribute to the organization of relational memory representations. The presence of one versus two categories did not alter either overall relational memory accuracy or the number of overall misassignments. However, on trials when there were two categories present, visually similar stimuli (i.e., stimuli from the same color [within-category, or WC] category) were more likely to be misassigned to each other's locations than was the case for those from different color (between-category, or BC) categories. This was true despite the fact that there were more opportunities to make BCM (i.e., 18 opportunities per trial) compared to WCM (i.e., 12 opportunities per trial). Thus, these data suggest categorical information was bound into the memory representation and that the presence of multiple categories of stimuli influenced the organization of the representation. It appears that in the presence of two categories of stimuli, memory for individual stimuli within a category was “fuzzier” leading to predictable WC memory errors. Interestingly, the presence of two categories did not impact overall memory performance in this experiment, only the nature of the memory errors committed.
Experiment 2
Experiment 1 demonstrated that perceptual (i.e., color) categories influenced the organization of the memory representation that then affected the fidelity of spatial relational memory. In Experiment 2, we extended these findings to investigate whether semantic categories also influenced the organization of relational memory representations. In this experiment, on each trial participants studied eight nameable objects either from a single semantic category (e.g., birds; one-category trials) or two semantic categories (i.e., birds and musical instruments; two-category trials; Fig. 3). Eight categories were used throughout the experiment (see Materials and Methods): mammals, kitchenware, fruits, tools, birds, musical instruments, vegetables, and furniture.
Example Experiment 2 study trial display including stimuli from two semantic categories (musical instruments and birds).
Average relational memory accuracy was 59.9% (SD = 5.6%) on one-category trials and 55.1% (SD = 5.6%) on two-category trials; the difference was significant, t(49) = 5.63, P < .001, d = 0.80. The average misplacement error was 143.5 pixels (SD = 45.3) on one-category trials and 150.9 pixels (SD = 41.5.8) on two-category trials. The difference was significant, t(49) = −2.65, p = 0.011, d = 0.37. The average number of misassignments (Fig. 4) was 7.2 (SD = 6.5) for one-category trials and 8.9 (SD = 6.8) for two-category trials and this difference was significant, t(49) = −2.90, p = 0.006, d = 0.42. For two-category trials, the average total number of WCM was 5.5 (SD = 5.0), and the average total number of BCM was 3.4 (SD = 2.8). The difference between WC and BC was significant, t(49) = 3.45, P < 0.001, d = 0.48.
Experiment 2 misassignment data. Total number of misassignment errors for one-category and two-category trials (blue bars) and misassignment errors for two-category trials separated into within-category misassignment (WCM) errors and between-category misassignment (BCM) errors (green bars). Plotted error bars indicate SEM. (***) P < 0.001.
Experiment 2 showed that semantic categories also influenced the fidelity of spatial relational memory. In a conceptual replication of Experiment 1, when there were multiple categories present on a given trial, participants were more likely to misassign an object to a location that had been occupied by another object from the same semantic category than the other semantic category. All of this despite the fact that the stimuli were all distinct, nameable objects and participants made, overall, very few misassignment errors (i.e., half as many as Exp 1 despite more objects per trial). In contrast to Experiment 1, participants were significantly less accurate and made significantly more misassignment errors on two-category trials compared to one-category trials. Given that on two-category trials, participants were more likely to make WCM than BCM errors, one might expect that on one-category trials, when all objects were from the same semantic category, participants would find objects similarly confusable and make even more errors. However, the opposite was true, suggesting that the presence of stimuli from multiple categories changed how the display was organized in memory, resulting in poorer performance on two-category trials. Together, these data again suggest that the presence of multiple semantic categories can contribute to the organization of relational memory representation and influence the patterns of errors.
Experiment 3
Experiment 2 extended the findings from Experiment 1 demonstrating that not only perceptual, but also semantic category information is bound into the memory representation for each study trial influencing its organization, which in turn influences behavioral outcomes. An important caveat, however, is that items within a semantic category are more perceptually similar to each other than to items from a different semantic category. Indeed, robins look more like eagles than they do like trombones. In Experiment 3, we sought to control visual similarity and distinctiveness across stimuli by using morph-line stimuli that were systematically varied such that adjacent morphs were equally perceptually similar to each other, but fell into two distinct semantic categories (i.e., cats and dogs; Fig. 5). The goal of Experiment 3 was to determine whether semantic categories still produced more WCM than BCM errors when controlling for perceptual similarity. In addition to the spatial reconstruction task, Experiment 3 also included a stimulus familiarization task to ensure that participants could differentiate between more dog-like and more cat-like stimuli. Briefly, in that task, participants were first shown a centrally presented morph with the cat and dog prototypes presented on either side. Participants indicated if that morph was more like the cat or the dog. Next, morph stimuli were shown alone (i.e., without the prototypes) and participants again indicated if that morph was more cat-like or more dog-like. Task accuracy was calculated. Task order (i.e., stimulus familiarization task vs. spatial reconstruction task) was counterbalanced across participants.
Sample morph line stimuli. Six stimuli from one of the cat-dog morph lines used in Experiment 3. Each step in the morph line is altered 20% in each dimension (cat vs. dog).
In the stimulus familiarization task, a 2 (task type: prototype present vs. prototype absent) × 2 (task order) repeated measures ANOVA revealed a significant main effect of task type, F(1,45) = 13.19, P =0 .001; neither the main effect of task order, F(1,45) = 0.82, P = 0.371, nor the interaction, F(1,45) = 0.57, P = 0.453, were significant. The task order was collapsed for subsequent analyses. Participants were 90.1% (SD = 13.8%) accurate at classifying the stimuli with the prototype stimuli present and 88.4% (SD = 13.9%) accurate without the prototype stimuli, and the difference was significant, t(46) = 3.63, p = 0.001, d = 0.53.
In the spatial reconstruction task, average relational memory accuracy was 36.7% (SD = 12.8%). For the purpose of conducting WC and BC comparisons, stimuli on each trial were categorized as “dog” if they were the three most dog-like morphs (i.e., D1, D2, and D3) and “cat” if they were the three most cat-like morphs (i.e., C3, C2, and C1). A 2 (error type: WCM vs. BCM) × 2 (task order: spatial reconstruction vs. stimulus familiarization) ANOVA revealed a significant main effect of error type, F(1,47) = 27.6, P < 0.001; neither the main effect of task order, F(2,47) = 0.954, P = 0.393, nor the interaction, F(2,47) = 2.14, P = 0.129, was significant. Therefore, task order was collapsed in all subsequent analyses. The average total number of WCM was 25.3 (SD = 11.8) and the average total number of BCM was 8.6 (SD = 6.1). The difference between WC and BC was significant, t(49) = 10.0, p < 0.001, d = 1.4.
An advantage of the morph stimuli is that they are ordered based on visual similarity to each other and the number of misassignments can be computed for each stimulus pair (Table 1). Indeed, while these stimuli can be dichotomized as more cat-like and more dog-like, adjacent stimuli are always equally visually similar to each other across the morph line, regardless of assigned category membership. Thus, D1 and D3 are both dog-like stimuli, but D3 is actually more visually similar to C3 than it is to D1. To assess the relationship between stimulus similarity and category membership, we next analyzed data across the full morph line. Unsurprisingly, as seen in Table 1, the most frequent misassignments occurred between stimuli that were adjacent to each other along the morph line. However, not all adjacent stimuli were misassigned with the same frequency. Indeed, a repeated measures 5 × 1 ANOVA including misassignments of adjacent stimuli (in the morph line, not on the screen; Fig. 6) revealed a significant main effect, F(3.6,174.9) = 12.6, P < 0.001. Follow-up post hoc t-tests showed that participants made significantly more misassignments between C1 and C2 compared to C2 and C3, t(49) = 2.5, P = 0.016, d = 0.41, and C2 and C3 compared to C3 and D3, t(49) = 2.8, P = 0.007, d = 0.40. Similarly, participants made significantly more misassignments between D1 and D2 compared to D2 and D3, t(49) = − 3.0, p = 0.004, d = 0.42, and significantly more misassignments between D2 and D3 compared to D3 and C3, t(49) = − 3.5, p = 0.001, d = 0.49. In short, more misassignments were made for more prototypical members WC, while fewer misassignments were made for morphs that were at the category boundary, even though visual similarity was equal between these stimuli. These data suggest a semantic boundary was imposed such that participants were least likely to misassign adjacent stimuli from different semantic categories.
Distribution of misplacement errors for neighboring morphs in Experiment 3. Total number of misassignment errors for each of the neighboring morph pairs (see Fig. 5): Cat1 versus Cat2; Cat2 versus Cat3; Cat3 versus Dog3; Dog3 versus Dog2; and Dog2 versus Dog1. Plotted error bars indicate SEM. (*) P < 0.05, (**) P < 0.01.
Total misassignments for each morph pairing
The results from Experiment 3 replicate within/between-category error findings from Experiment 2, strengthening the finding that semantic categorical information is used to organize relational memory representations and influence task performance. Important to note, these data do not suggest that semantic categories are more salient than perceptual categories for organizing memory. They simply indicate that when perceptual similarity is carefully controlled, semantic categories are important for the organization of relational memory representations and also that the increase in WCM errors in Experiment 2 was not due only to perceptual similarity of category members. Furthermore, the misassignment rate was highest for category members that were the most similar to prototypical category exemplars (i.e., C1–C2 and D1–D2), consistent with findings indicating that individuals rely heavily on prototype information when making category judgments for new stimuli (Bowman and Zeithamova 2018, 2023).
Discussion
In three experiments, we demonstrated that the inclusion of categorical information can influence how spatial information is remembered. When participants were asked to remember the spatial location of multiple stimuli and those stimuli all either came from the same category or two distinct categories, participants either made an equal number of errors (Experiment 1) or made more errors on two-category trials (Experiment 2). To some, these findings may be somewhat surprising as one might expect that having more distinct stimuli would lead to a more precise memory representation (e.g., Alvarez and Cavanagh 2004; Awh et al. 2007). However, these data do not support this reasonable expectation. Rather, the data suggest that when two categories are present, that categorical information influences how the episode is represented in memory resulting in a disproportionate increase in item-location binding errors among items from the same category. Indeed, across all experiments, on two-category trials participants consistently made significantly more WC misplacement errors than BC misplacement errors when categories were perceptually or semantically defined. These data suggest that participants may be over-reliant on gist memory (i.e., something blue went here) to organize their mnemonic representations at the detriment of focused verbatim memory (i.e., this specific blue shape went here) leading to WCM. This hypothesis should be evaluated systematically in future work.
These findings, however, are consistent with and complement and extend much of the context boundary literature, albeit in a very different paradigm and a different dimension (i.e., space vs. time). In the context boundary literature, it has been repeatedly demonstrated that given a stream of information, the presence of context boundaries can influence proximity judgments (DuBrow and Davachi 2013; Ezzyat and Davachi 2014; Davachi and DuBrow 2015) and event boundaries can have both a faciliatory and disruptive effect on memory binding (DuBrow and Davachi 2013; Gold et al. 2017). For example, while watching a television show, commercials inserted at an event boundary were remembered better than those inserted within an event (Boltz 1992). Conversely, when presented sequentially, items are remembered as being closer together in time when they shared a context compared to when they were associated with a different contexts, even when those stimuli were, in fact, temporally equidistant (Ezzyat and Davachi 2014; Davachi and DuBrow 2015). In the current paradigm, all necessary information is presented simultaneously rather than acquired over time, and there are not temporal context shifts, but rather multiple categories of stimuli are present such that categorical boundaries exist within the display. And still, a similar phenomenon is reported in Experiment 3, where stimulus similarity was juxtaposed to category membership: Items that shared a category were remembered as more similar (i.e., more likely to produce a misassignment error), and items at a category boundary were represented as less similar (i.e., less likely to be misassigned), even though the visual similarity between adjacent pairs was the same in both situations.
Interestingly, this category version of the spatial reconstruction task shares many similar features to classic visual short-term memory paradigms with interesting parallels related to task outcomes. Studies of visual short-term memory tasks have typically used either change-detection type tasks or continuous report cued-recall type tasks (for review, see Brady et al. 2011). In a change detection task, a display of multiple items (e.g., colored squares, bars at various orientations, Chinese characters, etc.) is briefly presented and after a very short delay, a new display is presented, and participants must decide if the new display is the same or different from the display that was studied moments ago (e.g., Pashler 1988; Luck and Vogel 1997; Gilchrist and Cowan 2014). Memory performance is diminished for displays of complex stimuli (e.g., three-dimensional cubes or Chinese characters) compared to displays of simple stimuli (e.g., color patches). It has been suggested that complex stimuli are remembered with lower resolution and are thus less discriminable resulting in memory errors (e.g., Alvarez and Cavanagh 2004; Eng et al. 2005; Awh et al. 2007) resulting in comparison errors (e.g., Scolari et al. 2008; Fukuda et al. 2010) and poor memory outcomes. These data are, in many ways, consistent with the current set of experiments where the item-location bindings for more similar stimuli (i.e., stimuli from the same category) are more error prone than item-location bindings for more distinct stimuli (i.e., stimuli from different categories). Through this lens, it is plausible that the reason that WCM occurs with greater frequency than BCM in the spatial reconstruction task is that WC items are less distinctive and remembered with poorer resolution resulting in greater confusability and more comparison errors between items in memory vs. items on the screen at reconstruction time. Furthermore, as indicated in our Experiment 3, category membership can bias memory even for objectively similar items (e.g., Cat2 and Cat3 vs. Cat3 and Dog3). Future work should empirically investigate this possibility.
Perhaps even more relevant are findings from studies using continuous report-cued recall tasks. In continuous report cued-recall tasks, a display of multiple items (e.g., color patches, objects, etc.) is presented briefly and after a short delay, a single item from the study display is cued and participants must make a determination (e.g., location) about that item with a free response (i.e., indicating a specific location on the screen; Wilken and Ma 2004; Zhang and Luck 2008). When asked to indicate what location an item had been studied in, participants often make errors and when an error is made, it is typical for the indicated location to be the location of a similar stimulus, rather than a dissimilar stimulus (Schneegans and Bays 2017; Pratte 2019; Markov and Utochkin 2022). Indeed, Markov and Utochkin (2022) showed that when a study array included four items, two from each of two categories (e.g., apples and toy soldiers), participants were far more likely to erroneously put a test item into the location of the other item from the same category rather than the location of an item from the other category. Chen and Cowan (2013) have suggested that participants engage in a minimal-responder model in this task, such that only remembered information about the single tested stimulus is considered when a response is selected, and if no information is available, participants randomly guess. Indeed, random guessing is typical in this task (Adam et al. 2017). While this pattern of data certainly mirrors the findings from the current study where WCM were more common than BCM, it is, we believe, unlikely that participants use the minimal-responder model to shape their behavior.
There are two main reasons that random guessing seems unlikely in the current study. First, in the spatial reconstruction task, stimulus positions were unrestrained such that stimuli could appear anywhere in the screen. Across all three experiments, the stimuli only occupy between 27° × 27° and 35° × 35° of visual angle, so the majority of the screen was unoccupied. As such, random guessing would be most likely to produce an error where the item was placed in a totally unstudied space. Yet, this is not what the current data indicate. Indeed, participants placed stimuli in 73.3% of studied locations (i.e., accurate placements and misassigned placements) in Experiment 1, 63.8% of studied locations in Experiment 2, and 68.3% of studied locations in Experiment 3. Second, unlike standard continuous report cued-recall tasks where one stimulus is tested at a time (see Adam et al. 2017 for an exception), in the spatial reconstruction task, all stimuli are tested simultaneously, so it is likely that participants are able to use a more process-of-elimination strategy for items that they are unsure about consistent with an ideal-responder model. This assumes that individuals remember the overall shape of the display and have a sense for about/generally where some stimulus belongs. Indeed, Horecka et al. (2018) demonstrated that even individuals with dense amnesia were able to remember the Gestalten shape-like shape of the display in this task, so this approach is plausible for the neurologically intact individuals who participated in the current study. Future work needs to empirically evaluate the possibility of minimal-responder versus ideal-responder model approaches to this task, but the current data, we believe, are suggestive.
While the current set of experiments provides interesting preliminary evidence that the inclusion of multiple stimulus categories can influence the organization of relational memory representations, resulting in predictable memory errors, considerably more research is required to understand if, for example, all category types produce a similar organizational effect or if this is specific to perceptually and semantically defined categories. Future studies could also explore the role of attention on memory outcomes in this task, perhaps by using ambiguous stimuli that could fall into different categories depending on task instructions; for example, stimuli that vary along two dimensions, one of which is task-relevant. Future work should also investigate the role of proactive interference that is known to influence memory performance when stimuli from the same category are used repeatedly across trials (Wickens 1970; Hubbard et al. 2018). The addition of eye tracking to future investigations should also be considered to better understand how study-time viewing behavior and/or viewing strategies can influence later memory outcomes (Lucas et al. 2019). Finally, while performance on the standard (noncategorical) spatial reconstruction task has been shown to depend critically on the hippocampus (Watson et al. 2013; Schwarb et al. 2016; Horecka et al. 2018), it is likely that performance on the categorical version depends on the broader hippocampal-memory dependent network (Ritchey et al. 2015; Wang et al. 2015; Rubin et al. 2017; Dulas et al. 2021), and elucidating the unique contributions of these structures will likely importantly contribute to our understanding of how memory representations are organized to affect memory outcomes.
Together, these data add to an already substantial, but still expanding, body of research highlighting the constructive nature of human memory. The current series of experiments demonstrated that multiple streams of information (i.e., item information, spatial information, and category information) interact to organize memory, and that the information bound on a given trial became the organizing principle that then influenced memory outcomes. Categorical information embedded in the modified spatial reconstruction task appeared to be used by participants to organize information in memory (i.e., chunking), even though it was neither necessary to achieve the task goals (i.e., remembering where an individual item was presented in space) nor without consequences (i.e., producing increased misplacement errors). These data contribute to our understanding of how memory is organized through the intentional and automatic binding of the information present in a given scene or episode.
Materials and Methods
General
Participants
All participants were undergraduate students recruited from the University of Illinois at Urbana-Champaign. Study procedures were approved by the University of Illinois at Urbana-Champaign Institutional Review Board, and all participants were treated according to the guidelines approved by the American Psychological Association (American Psychological Association 2017). Participants were paid $8 for their participation.
Sample size was determined a priori based on a pilot study (N = 15) with an identical design to the design used in Experiment 1 comparing WC and BC misplacements (described in detail below). G*Power 3.1 (Faul et al. 2007) was used to calculate the sample size based on the pilot effect size of 0.56 with a significance criterion of α = 0.05 and power = 0.95. The power analysis identified a sample size of 44 participants per experiment. As such, the obtained sample size of N = 50 is more than adequate to investigate the study question.
Stimuli and apparatus
Stimuli were presented on a 21 in monitor using Presentation software (Neurobehavioral Systems). Participants were positioned in a chin rest 60 cm from the screen.
Experiment 1
Participants
Fifty-one naïve participants (ages 18–30; nine males) participated in Experiment 1. Data for one participant were corrupted and those data were excluded resulting in the target sample size of 50 participants.
Stimuli
Stimuli included abstract line drawings that were filled-in in either blue or yellow. There were 60 blue stimuli and 60 yellow stimuli. As such, stimuli were sorted into two distinct color categories. Each stimulus was only used once during the experiment. Each stimulus was 4.5° × 4.5° of visual angle. Participants were eye-tracked, and those data are presented elsewhere (Lucas et al. 2019).
Design and procedure
Participants were given the following instructions verbally: “In this study, you will be looking at a series of displays on the computer screen like this one (example trial printed on a piece of paper), each containing six abstract objects. As you can see, some of the objects are blue and some are yellow. Sometimes you will see yellow and blue objects, other times you will see just yellow or just blue objects. I want you to study each display and try to remember the locations of all six objects so that you can reproduce the display later. You'll be allowed to study each display for 16 sec and then the objects will disappear briefly. After a few seconds, the objects will reappear at the top of the screen and your job is to put each object back in the location in which it was studied. The goal is to reconstruct the study display as accurately as possible. Let's try a few practice trials together.” Each participant then completed six practice trials before the start of the experimental task. Practice trials and task trials were identical.
On each study trial, six items were presented at six pseudo-random, nonoverlapping positions on the screen. X and y coordinates were randomly generated for any given stimulus in a display; and then if there was any resulting overlap among stimuli, new coordinates were randomly drawn until there were no overlapping stimuli in a display. On two-category trials, the first three x–y coordinate pairs drawn were assigned to category A, and the last three x–y coordinate pairs were assigned to category B. The study display remained on the screen for 16 sec. All items were then removed from the screen for a 4 sec delay. Following the delay, the items reappeared in a straight line at the top of the screen. Participants then used the mouse to move each item back to its remembered location. Items could be moved around until the participant was satisfied with their reconstruction; time was unlimited. Participants indicated that they had completed their reconstruction by pressing the space bar. The next trial began 2000 msec after the space bar was pressed. There was a total of 40 trials; 20 were single-category trials, 20 were dual-category trials. Participants were always shown alternating blocks of 10 single-category and 10 dual-category trials. The presentation order was counterbalanced.
Experiment 2
Participants
Fifty-two naïve participants (ages 18–39; 17 males) participated in this experiment. Due to significant skewedness in the data, median absolute deviation methods were used to detect statistical outliers at the conservative threshold of 3 (Hampel 1974; Leys et al. 2013), and two participants were excluded from the analysis. The resulting sample size was 50 participants.
Stimuli
Stimuli included black and white images of real-world objects. Objects were drawn from the Bank of Standardized Stimuli (BOSS) Normative Photos database (Brodeur et al. 2014) that have been normed for category membership. Category agreement scores were > 75% for all selected categories, and category pairs were selected to always include one living and one nonliving category to ensure maximum dissimilarity between categories. Objects were converted to grayscale using Photoshop. The objects fell into eight categories: mammals, kitchenware, fruits, tools, birds, musical instruments, vegetables, and furniture. Each stimulus was only used once during the experiment. Each stimulus was 4.4° × 4.4° of visual angle.
Design and procedure
Participants were given the following instructions verbally: “In this study, you will be looking at a series of displays on the computer screen like this one (example trial printed on a piece of paper), each containing eight objects. As you can see in this example, some of the objects are foods and some objects are articles of clothing. Sometimes you will see food and clothes, other times, you will see just food or just clothes. I want you to study each display and try to remember the locations of all eight objects so that you can reproduce the display later. You'll be allowed to study each display for 24 sec and then the objects will disappear briefly. After a few seconds, the objects will reappear at the top of the screen and your job is to put each object back in the location in which it was studied. The goal is to reconstruct the study display as accurately as possible. Let's try a few practice trials together.” Participants then completed three practice trials before the start of the experimental task. Practice trials were identical to task trials, except that the stimuli came from the categories clothing and foods (non-fruits and vegetables).
The design was identical to Experiment 1 with three exceptions. First, there were four blocks of eight trials each (four single-category and four dual-category trials). Second, on each trial, eight stimuli were presented; on two-category trials, there were four stimuli from each category. The number of stimuli per trial was increased from six abstract shapes in Experiment 1 to eight nameable objects in Experiment 2, based on pilot data to better equate difficulty across the experiments. The category pairings were mammals and kitchenware, fruits and tools, birds and musical instruments, and vegetables and furniture; one pairing per block. Category pairings were chosen to maximize BC semantic dissimilarity by always pairing a living category with a nonliving category. Third, the study phase was extended to 24 sec to accommodate the additional study stimuli.
Experiment 3
Participants
Fifty-one naïve participants (ages 18–22; 15 males) from the University of Illinois at Urbana-Champaign participated in this experiment. Data for one participant were corrupted and could not be analyzed, resulting in a sample of 50 participants.
Stimuli
Stimuli included cat-dog animal morphs. Morphs were generated by applying an algorithm (Shelton 2000) to three cat and three dog prototype images. Morphs were created such that there were four intervening morphs between each prototype. For example, if D1 and C1 were the prototype, the four intervening morphs would be: (1) 80% D1, 20% D1; (2) 60% D1, 40% C1; (3) 40% D1, 60% C1; and (4) 20% D1, 80% C1 (Fig. 5). Every combination of cat and dog prototypes was included resulting in nine distinct stimulus sets, each with six stimuli. Each set was used twice in this experiment for a total of 18 trials. Each stimulus was 4.6° × 4.6° of visual angle. These morphs have been used in previous studies (Shelton 2000; Hassevoort et al. 2018).
Design and procedure
There were two separate tasks in Experiment 3: A spatial reconstruction task using the morph stimuli and a stimulus familiarization task. The purpose of the stimulus familiarization task (described below) was to familiarize participants with the unusual morph stimulus set and train them to distinguish cat stimuli from dog stimuli. The idea was that having familiarity with categorizing the stimuli may enhance the category effect on the spatial reconstruction task. There were, however, no significant differences on spatial reconstruction task performance for those who completed the familiarization task first, so data were collapsed for the final analyses.
The stimulus familiarization task was divided into two phases. During the first phase, a single morph stimulus was centrally presented, and the cat and dog prototypes were presented on either side (counterbalanced across participants). On each trial, the participant indicated if the centrally presented morph was more like the cat prototype or more like the dog prototype. Feedback was provided after each trial, and there were 108 such trials. The second phase was similar to the first except that the prototypes were no longer provided. Again, feedback was provided after each trial, and there were a total of 108 trials. Data were overwritten for three participants and could not be included in the analyses.
For the spatial reconstruction task, participants were given the following instructions verbally: “In this part of the study, you will be looking at a series of displays on the computer screen, each containing six cat-dogs in different locations on the screen. Your goal when studying each display will be to try to memorize the locations of all six cat-dogs so that you can reproduce the display later. You'll be allowed to view each display for about 16 sec. Shortly after you view each display, you will be asked to reconstruct the display as accurately as possible by moving each object back to its original location. Let's start with three practice trials.” Participants then completed three practice trials. Practice trials were identical to task trials, except that the morph stimuli were not used; instead, stimuli were six squares labeled CAT1, CAT2, CAT3, DOG1, DOG2, DOG3.
The design of the spatial reconstruction task was identical to Experiment 1, except that there was a single block with 18 trials. The order of the spatial reconstruction task and the stimulus familiarization task was counterbalanced across participants.
Data processing and statistical analysis
The spatial reconstruction task analysis has been previously described in detail (Horecka et al. 2018) and was applied here. The data processing pipeline is also available through gitbhub (https://github.com/kevroy314/msl-iposition-pipeline). Generally, data were analyzed by comparing the location of each stimulus in the studied display to the location of each stimulus in the reconstructed display. This involved four steps. First, information about the identity of each stimulus was ignored, resulting in one set of six x–y coordinates per trial for the studied display and another set of six x–y coordinates per trial for the reconstructed display. Next, global errors, those systematic spatial errors that were shared across all items in the reconstructed display (i.e., global rotation, global scaling, and global translation), were corrected. Then, identity information was reintroduced to determine whether each item of the reconstruction was assigned to its own studied location, another item's studied location, or an unstudied location.
Overall relational memory accuracy was computed by identifying the number of items placed in their studied location divided by the total number of items (i.e., six in Experiments 1 and 3; eight in Experiment 2). Accuracy was computed separately for one- and two-category trials in Experiments 1 and 2.
Misplacement errors were calculated by computing the distance (in pixels) between where an individual item was placed at test and where that item had been studied for one- and two-category trials separately in Experiments 1 and 2.
Misassignments were computed by determining the number of items that were placed back into a studied location, but not the correct location where that specific item had been studied. In other words, misassignments occurred when an incorrect item was placed back into one of the studied locations. On two-category trials, these misassignments were then sorted based on whether a misassigned item was assigned to the studied location of an item in its same category or an item from the other category. Note that in Experiment 3, stimuli that were more than 50% cat were categorized as “cats” and stimuli that were more than 50% dog were categorized as dogs to create two categories: cats and dogs. Two-category trial performance was further separated into WCM and BCM. For Experiments 1 and 2, paired sample t-tests were performed (SPSS version 29) comparing accuracy, misplacement errors, and misassignment rates across trial types; Cohen's d is reported (G*Power 3.1).
For Experiment 3, to assess whether task order affected performance, a 2 (error type: WCM vs. BCM) × 2 (task order) ANOVA was computed. As in Experiments 1 and 2, paired sample t-tests were performed comparing misassignment rates across trial types (WCM vs. BCM). Misassignments for each morph pair were also computed. A repeated measures ANOVA was performed on adjacent pairs (e.g., C1 and C2, C3 and C3, C3 and D3, etc.), and post hoc paired sample t-tests were conducted to assess differences. Huynh–Feldt adjustment and Cohen's d are reported throughout where appropriate.
Acknowledgments
This work was funded by National Institute of Mental Health grant R01MH062500 awarded to Neal J. Cohen.
Footnotes
-
Article is online at http://www.learnmem.org/cgi/doi/10.1101/lm.054006.124.
- Received March 11, 2024.
- Revision received September 24, 2024.
- Accepted October 1, 2024.
This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the first 12 months after the full-issue publication date (see http://learnmem.cshlp.org/site/misc/terms.xhtml). After 12 months, it is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.
















