A reward effect on memory retention, consolidation, and generalization?

  1. Roland G. Benoit1,3,4
  1. 1Max Planck Research Group: Adaptive Memory, Max Planck Institute for Human Cognitive and Brain Sciences, 04103 Leipzig, Germany
  2. 2International Max Planck Research School NeuroCom, 04103 Leipzig, Germany
  3. 3Department of Psychology and Neuroscience, University of Colorado Boulder, Boulder, Colorado 80309, USA
  4. 4Institute of Cognitive Science, University of Colorado Boulder, Boulder, Colorado 80309, USA
  1. Corresponding authors: heidrun.schultz{at}cbs.mpg.de, roland.benoit{at}colorado.edu
  1. 5 These authors contributed equally to this work.

Abstract

Reward improves memory through both encoding and consolidation processes. In this preregistered study, we tested whether reward effects on memory generalize from high-rewarded items to low-rewarded but episodically related items. Fifty-nine human volunteers incidentally encoded associations between unique objects and repeated scenes. Some scenes typically yielded high reward, whereas others typically yielded low reward. Memory was tested immediately after encoding (n = 29) or the next day (n = 30). Overall, reward had only a limited influence on memory. It did not enhance consolidation and its effect did not generalize to episodically related stimuli. We thus contribute to understanding the boundary conditions of reward effects on memory.

Each day, we experience a steady stream of episodes, the majority of which we will forget. What determines whether an event will be committed to long-term memory? One important factor is reward. The presence or anticipation of reward during encoding enhances memory. This is the case when the receipt of the reward is conditional on later success in remembering the memory (Adcock et al. 2006; Wolosin et al. 2012, 2013; Murty et al. 2017), but also when it is not (Wittmann et al. 2005, 2008; Krebs et al. 2009; Mather and Schoeke 2011; Bunzeck et al. 2012; Murty and Adcock 2014; Gruber et al. 2016; Schultz et al. 2022).

Reward augments memory formation already at its earliest stage. During encoding, it has been associated with coordinated activity of the hippocampus and the dopaminergic midbrain (Wittmann et al. 2005, 2008; Adcock et al. 2006; Wolosin et al. 2012). However, the reward effect further relies on postencoding processes. Specifically, during postencoding rest (Gruber et al. 2016) and sleep (Sterpenich et al. 2021), neural representations of rewarded memories are preferentially reactivated. Accordingly, the reward effect further increases after a consolidation period (Wittmann et al. 2005; Murayama and Kuhbandner 2011; Murayama and Kitagami 2014; Spaniol et al. 2014; Patil et al. 2017; but see Jang et al. 2019; Gholston et al. 2023).

Consolidation, however, does not merely strengthen individual episodic memories. It further results in more generalized memory representations that encode commonalities across related experiences (Tompary and Davachi 2017). Importantly, such generalization also extends the mnemonic benefit of reward to semantically related memories. For example, rewarding objects also enhances memory for unrewarded objects from the same semantic category—but only after consolidation (Oyarzún et al. 2016; Patil et al. 2017). Similar generalization occurs following aversive stimulation (Dunsmoor et al. 2015; Starita et al. 2019). Here, we hypothesize that reward also benefits memories that are not semantically but episodically related; for example, via shared associative features (see also Wimmer et al. 2014; Benoit et al. 2019; Palombo et al. 2021; Paulus et al. 2022).

We examined this hypothesis in a preregistered study (https://osf.io/35nbs/?view_only=acfb2cd5cc1a470596db82b87807be1e). Participants (n = 59) (see the Supplemental Material for details) performed an incidental memory-encoding task in which they learned to associate objects with background scenes (Gruber et al. 2016). As within-participant factors, we manipulated the actual reward magnitude on each trial (high or low) and whether the magnitude was typical for that scene (typical or atypical). As a between-subjects factor, we tested memory either immediately after encoding (immediate group; n = 29) or after 24 h (delayed group; n = 30). We predicted (1) that high reward enhances associative memory and (2) that the benefit of high reward generalizes from typical high-reward trials to atypical low-reward trials via their shared scene association. Both effects should be more pronounced after consolidation.

The procedure consisted of two parts (Fig. 1A). In part 1, participants performed the incidental encoding task (Fig. 1B). Here, they encountered 160 objects paired with one of four scenes. Participants mentally simulated a scene-specific action and responded to a corresponding question (e.g., “Would the object float on water?” for the swimming pool scene) (see the Supplemental Material). Each pair was presented twice. Correct responses to the question yielded a reward.

Figure 1.

(A) Experiment structure. In part 1, participants completed the incidental encoding task. After either 15 min (immediate group) or 24 h (delayed group), they completed part 2, consisting of the scene recall task and reward recall task. (B) Example trial for the incidental encoding task. After a variable intertrial interval (1.75–5.75 sec), each trial started with the presentation of a colored rectangle and a scene (0.75 sec), before the object appeared within the rectangle (3 sec). Participants were asked to mentally imagine a scene-specific action before responding to a corresponding question (1.5 sec). For correct responses, they received visual feedback of the earned reward magnitude (high or low; 1 sec). (C) Example trial for the scene recall task. In each trial, after a variable intertrial interval (1–5 sec), one of the objects was presented (3 sec), and participants either selected its paired scene or indicated that they had forgotten the scene (2 sec). They then rated their confidence on a four-point scale (1 [not at all] to 4 [very confident]; 2 sec). (D) Example trial for the reward recall task. In each self-paced trial, one of the objects was presented, and participants indicated whether it had been associated with a high or low reward or responded that they had forgotten the reward. For copyright reasons, we here display photographs that are similar to the actual experimental stimuli (scene image from https://pixabay.com; object image by the investigators).

Two scenes were typically associated with high rewards (typical high reward, 2 points), and two were associated with low rewards (typical low reward, 0.02 points). Points were later converted to monetary rewards (see the Supplemental Material). On 20% of the trials, a scene typically paired with a low reward yielded a high reward (atypical high reward) and vice versa (atypical low reward). The actual reward on each trial was indicated by the color of a rectangle superimposed on the scene (see the Supplemental Material). Color–reward associations were learned to criterion (100%) in the beginning of part 1 and tested again immediately after encoding (93%–98% accurate on average) (see Supplemental Table S8).

Part 2 started with the critical surprise scene recall task (Fig. 1C). For each object, participants were asked to indicate the paired scene from the encoding phase and to rate their confidence. Finally, in the reward recall task (Fig. 1D), they were asked to indicate whether each object had been paired with a high or low reward. Tasks were implemented in Matlab and PsychToolBox. However, due to a bug in the experiment code, the square object images in all tasks were stretched to a 4:3 format. Additionally, participants completed questionnaires (see the Supplemental Material).

We used generalized linear mixed models (GLMMs) in R (afex::mixed; https://CRAN.R-project.org/package=afex) to analyze response times (RTs) in the encoding task as well as accuracies in the scene recall task. For the encoding RTs, the model included the fixed effects of reward and typicality. For scene recall, one model included the fixed effects of time, reward, and typicality. Another model tested whether the reward effect on scene recall depended on participants’ ability to recall the associated reward. Hence, it included the additional fixed effects of reward recall. All GLMMs modeled the main effects and interactions of the fixed effects, as well as random effects (Supplemental Material; see Supplemental Table S1 for details on GLMMs and model selection; see Supplemental Tables S9–S11 for detailed output for each reported model). Data as well as R code for the GLMM analyses are available on OSF (https://osf.io/6vhgr). For further descriptive analyses, please see Supplemental Tables S2–S8.

Our analyses diverted from the preregistration as follows: Our preregistered analysis plan included separate GLMMs for typical and atypical trials. However, to statistically examine differences due to typicality, we first computed overall GLMMs with typicality as a fixed effect. As typicality did not interact with any other fixed effect, we only report the separate analyses for typical and atypical trials in Supplemental Tables S13–S22. They yielded largely the same result pattern as the analyses reported below.

First, we tested whether our reward manipulation succeeded in speeding up encoding RTs (see Fig. 2A; Supplemental Table S3), as in the similar experiment by Gruber et al. (2016). A GLMM with the fixed factors reward and typicality yielded this significant effect of reward (RTs high < low reward, F(1,96.434) = 6.245, P = 0.014) with no main effect of typicality or interaction (P ≥ 0.461) (see the Supplemental Material; Supplemental Table S9).

Figure 2.

Results. (A) Response times (RTs) in the incidental encoding task, shown separately for the reward and typicality conditions. Single points denote individual median RTs. (B) Proportions of correct scene recall as a function of time of test, reward, and typicality. (C) Proportions of correct scene recall, shown separately for items with correct and incorrect reward recall and aggregated over typicality. Note that due to missing values for some participants, sample size varies between conditions. Black frames indicate significant effects from the GLMMs (see the text for details). Inner error boxes denote the group-adjusted standard error of the mean, and outer error boxes denote the group-adjusted standard deviation (Cousineau 2005).

Second, we tested whether our experimental manipulations influenced memory for the associated scenes. Overall, scene recall rates were high (see Fig. 2B; Supplemental Table S4). A GLMM with the fixed factors reward, typicality, and time yielded a significant main effect of time (immediate > delayed, Χ2(1) = 11.492, P < 0.001). However, while scene recall was numerically higher for high-reward trials than for low-reward trials, the main effect of reward was not significant (Χ2(1) = 2.320, P = 0.128).

Furthermore, we found no evidence that reward influenced consolidation: Although the reward-related difference was numerically larger in the delayed group than in the immediate group, the interaction of time and reward was not significant (Χ2(1) = 1.614, P = 0.204). We also found no evidence for a reward effect on time-dependent generalization. We had expected consolidation to increase retention of low-rewarded memories containing scenes typically associated with high rewards (i.e., atypical low-reward trials); numerically, we found the opposite pattern (see Fig. 2B). However, the three-way interaction of time, reward, and typicality was not significant (Χ2(1) = 2.428, P = 0.119), and neither was any other main effect or interaction (P ≥ 0.218) (see the Supplemental Material; Supplemental Table S10).

Finally, we tested whether participants were better at recalling high-reward than low-reward memories when they also remembered the associated reward magnitude (see Fig. 2C; Supplemental Table S6). We computed a GLMM for scene recall with the additional fixed factor of reward recall. This yielded significant main effects of reward (high > low, Χ2(1) = 6.959, P = 0.008), time (immediate > delayed, Χ2(1) = 12.145, P < 0.001), and reward recall (correct > incorrect, Χ2(1) = 77.456, P < 0.001), as well as a significant interaction of reward and reward recall (Χ2(1) = 9.930, P = 0.002). No other effect was significant (P ≥ 0.196) (see the Supplemental Material; Supplemental Table S11).

Following up on the interaction, we computed paired comparisons of the high-reward and low-reward conditions separately for trials with correct and incorrect reward recall, pooling over typicality and time. The reward effect on scene recall was significant only when participants had also correctly recalled the reward (reward recalled: z ratio = 3.371, P < 0.001; not recalled: z ratio = 0.265, P = 0.791) (see Supplemental Table S12).

In sum, we found limited evidence for an overall reward effect on memory. Reward improved memory in a subset of trials for which participants also remembered the reward magnitude. The reward-related difference in scene recall was numerically larger in the Delayed group (see Fig. 2B,C), consistent with a role of consolidation. However, this interaction was not significant. Last, we saw no evidence for the reward effect to generalize to low-reward memories sharing the same scene association.

A positive effect of reward on memory formation is well established (Wittmann et al. 2005, 2008; Adcock et al. 2006; Krebs et al. 2009; Mather and Schoeke 2011; Bunzeck et al. 2012; Wolosin et al. 2012, 2013; Murty and Adcock 2014; Gruber et al. 2016; Murty et al. 2017; Schultz et al. 2022). Here, we adapted a paradigm that yielded a reward effect on immediate scene recall (Gruber et al. 2016). In that study, memory performance was overall low (approximately one out of four recalled trials). To avoid floor effects in the delayed condition, we thus presented each object–scene pair twice. This resulted in substantially higher recall rates across conditions (see Fig. 2B,C; Supplemental Table S4).

However, the repeated presentation may have contributed to the limited overall reward effect on memory. First, this effect is thought to be based on the mesolimbic reward system (Wittmann et al. 2005, 2008; Adcock et al. 2006; Wolosin et al. 2012; but see Gieske and Sommer 2022). The repetition may have attenuated the mesolimbic response by reducing both the reward prediction error (Schultz 1998; Rouhani et al. 2018; Rouhani and Niv 2021) and stimulus novelty (Bunzeck and Düzel 2006; Wittmann et al. 2007; Krebs et al. 2009; Bunzeck et al. 2012).

Second, reward may selectively boost memory following shallow compared with deep encoding (Shigemune et al. 2017; Yan et al. 2022); for example, through synaptic tagging (Frey and Morris 1997; Redondo and Morris 2011). In contrast, the repetitions in our task likely deepened encoding, potentially making memory formation less susceptible to reward.

The reward effect was numerically larger after a 1-d consolidation period, though the critical interaction was not significant. Previous work has largely demonstrated that consolidation boosts reward effects on memory (Wittmann et al. 2005; Murayama and Kuhbandner 2011; Murayama and Kitagami 2014; Spaniol et al. 2014; Patil et al. 2017; cf. Jang et al. 2019; Gholston et al. 2023). This may reflect the preferred reactivation of rewarded items during slow-wave sleep (Sterpenich et al. 2021).

However, there is evidence that neural replay prioritizes weakly encoded items (Schapiro et al. 2018; Denis et al. 2021). Thus, the repeated encoding in our study may have created strong memories that were less likely to be replayed overall. As a consequence, the consolidation period would have contributed less to the further strengthening of any memories, including the rewarded memories.

Consolidation can also promote the generalization of motivationally salient memories; that is, an initial mnemonic benefit for items directly associated with reward or aversive stimulation may generalize, after consolidation, to other items from the same semantic category (Dunsmoor et al. 2015; Oyarzún et al. 2016; Patil et al. 2017; Starita et al. 2019). Here, we examined whether such generalization can also occur when the items are episodically related.

Such generalization would have affected the reward effect for atypical trials in the delayed group: Over time, atypical low-reward trials would have received a memory benefit from the scene association that they shared with typical high-reward trials (i.e., delayed group: atypical low reward > typical low reward). We did not find such an effect. On the contrary, the atypical low-reward memories were remembered worst in the delayed group.

Why did we not see evidence for generalization? First, generalization may require pre-experimental (i.e., semantic) associations. Thus, it would not extend to items that are only episodically associated. However, there is evidence that neural representations of memories become more similar over time when they share episodic associations (Tompary and Davachi 2017). Moreover, value has been shown to spread along the edges of episodic associative networks (Wimmer et al. 2014; Benoit et al. 2019; Paulus et al. 2022).

Second, previous studies found generalization of the reward effect to nonrewarded stimuli of the same semantic category that were presented in different experimental blocks (Oyarzún et al. 2016; Patil et al. 2017). As participants had no reason to expect rewards in these blocks, they would have experienced small or no reward prediction errors (Schultz 1998). In contrast, we presented typical and atypical trials in an interspersed manner. If participants had reward expectations based on the shared scene association, atypical trials would yield larger prediction errors than typical trials (i.e., positive prediction errors for atypical high-reward trials and negative prediction errors for atypical low-reward trials). This, in turn, may have affected memory formation for the atypical trials (Rouhani et al. 2018; Rouhani and Niv 2021), potentially counteracting the generalization effect. It is unclear, however, whether such expectations played a role, given that (1) a majority of the participants did not notice the relationship between scenes and reward magnitudes (see the Supplemental Material) and (2) the color cues provided a more reliable (i.e., deterministic) predictor of the upcoming reward than the scenes.

The effect of reward on scene recall was present only when participants also remembered the reward. This does not necessarily mean that this effect hinges on explicit recall of the rewarding outcome. It could also suggest that the reward effect was selective for strongly remembered trials; that is, trials for which participants also remembered other associated information (such as the reward magnitude). We note that few participants explicitly noticed the scene–reward associations. Thus, they were unlikely to merely infer the reward magnitude from the recalled scene or vice versa. There is, however, some recent evidence that participants have a bias to assume that recollected memories had been associated with a high reward (Schultz et al. 2022). Such a bias may have influenced the present results.

Did our study have sufficient statistical power? We concluded data collection after 1 yr with a sample size of n = 59, falling one participant short of our stopping criterion of n = 60 after 1 yr. Our final sample size still yielded a power of 0.85 for the reward generalization effect reported by Patil et al. (2017) (see the Supplemental Material for details on sample size calculation). We note that due to the difficulty of calculating power for GLMMs with complex random effect structures (Kumle et al. 2021), our calculations are based on t-tests.

Could our choice of outcome explain why we did not see more pronounced effects? A number of studies have demonstrated reward effects on recognition memory (Wittmann et al. 2005; Adcock et al. 2006; Murty and Adcock 2014; Gruber et al. 2016; Schultz et al. 2022), including those reporting generalization to semantically related memories (Oyarzún et al. 2016; Patil et al. 2017). In contrast, we tested associative recall. There were two reasons for this: First, our paradigm necessarily deviated from previous semantic generalization paradigms in which participants encoded and recognized single objects of different semantic categories. Here, we investigated generalization along episodic associations. Thus, we used a deep encoding paradigm (Gruber et al. 2016; Tompary and Davachi 2017) to create integrated memory traces that comprised an object, a scene, and their association. The associative recall task allowed us to probe the entire memory (Horner et al. 2015). Second, reward has been shown to enhance hippocampal memory; that is, recollective, relational, or associative memory (Wittmann et al. 2005; Davachi 2006; Diana et al. 2007; Wolosin et al. 2012, 2013; Gruber et al. 2016; Eichenbaum 2017). Associative recall thus seems well suited for assessing reward effects on memory. Moreover, an additional power analysis yielded a power of β = 0.98 to replicate the reward effect on associative scene recall in Gruber et al. (2016) (see the Supplemental Material). It is therefore unlikely that our choice of outcome was responsible for the reported null effects.

We note that our paradigm is somewhat complex compared with similar studies. It is possible that this hindered participants in tracking reward probabilities throughout the task. We argue, however, that our task succeeded in eliciting reward expectations: Responses in the encoding task were faster on high-reward trials than low-reward trials, replicating Gruber et al. (2016). Furthermore, memory for both typical and atypical color–reward associations was near ceiling immediately after encoding. This implies that participants were able to track reward probabilities for typical versus atypical trials throughout the task.

As a caveat, we note that our study differed from previous work in two key aspects: stimulus repetition and inclusion of atypical trials, whose potential impact we discussed above. We therefore cannot conclude with certainty whether our null effects were due to differences in study design or whether indeed the effects of reward are less reliable than previously thought. However, given the strengths of our study (preregistration, a priori power calculation, and a strong empirical basis for the hypothesized effects), we believe that it helps elucidate possible boundary conditions of reward effects on episodic memory.

Going further in that direction, future work could tackle the following: First, one could conduct a similar experiment without trial repetitions. While these are undoubtedly beneficial (e.g., for improving memory accuracy in experiments with long retention intervals), we discussed above how they may have interfered with a possible reward effect. Second, rather than interspersing atypical trials, one could use a block design akin to Oyarzún et al. (2016) and Patil et al. (2017). Here, object–scene pairs may first be presented in a nonrewarded block, followed by a rewarded block pairing the same scenes with different objects. Such a design would mitigate the discussed possible confound of reward prediction errors. If the reward effect generalizes after consolidation, not only would high-reward object–scene pairs from the rewarded block be remembered better, but so would pairs from the nonrewarded block that share the same high-reward scene.

In sum, in this preregistered study, we found limited evidence for either a reward effect on memory, its strengthening with consolidation, or its generalization to memories that are episodically related. These results suggest that the impact of reward on memory may be more elusive than commonly reported in the literature. They therefore also highlight the need to better understand the boundary conditions under which reward effects on memory emerge.

Data Deposition

Data as well as R code for the GLMM analyses are available on OSF (https://osf.io/6vhgr). The preregistration is at https://osf.io/35nbs/?view_only=acfb2cd5cc1a470596db82b87807be1e.

Competing interest statement

The authors declare no competing interests.

Acknowledgments

We thank Matthias J. Gruber for sharing the experimental stimuli from Gruber et al. (2016); Johanna Fiebig, Sarah-Lena Schäfer, Martina Dietrich, Elina Williamson, and Nuno Busch for their assistance in data collection; and Nico Scherf for advice on GLMMs. This work was supported by a Max Planck Research Group awarded to R.G.B.

Footnotes

  • Received June 26, 2023.
  • Accepted August 9, 2023.

This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the first 12 months after the full-issue publication date (see http://learnmem.cshlp.org/site/misc/terms.xhtml). After 12 months, it is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.

References

| Table of Contents