Positive affect amplifies integration within episodic memories in the laboratory and the real world
- Julia G. Pratt1,
- Stephanie E. Wemm1,
- Bailey B. Harris2,
- Yuye Huang3,
- Rajita Sinha1 and
- Elizabeth V. Goldfarb1,4,5,6
- 1Department of Psychiatry, Yale University, New Haven, Connecticut 06511, USA
- 2Department of Psychology, University of California Los Angeles, Los Angeles, California 90095, USA
- 3Department of Psychology, Johns Hopkins University, Baltimore, Maryland 21218, USA
- 4Department of Psychology, Yale University, New Haven, Connecticut 06511, USA
- 5Wu Tsai Institute, Yale University, New Haven, Connecticut 06511, USA
- 6National Center for PTSD, U.S. Department of Veterans Affairs, West Haven, Connecticut 06477, USA
- Corresponding author: elizabeth.goldfarb{at}yale.edu
Abstract
Emotional events hold a privileged place in our memories, differing in accuracy and structure from memories for neutral experiences. Although much work has focused on the pronounced differences in memory for negative experiences, there is growing evidence that positive events may lead to more holistic, or integrated, memories. However, it is unclear whether these affect-driven changes in memory structure, which have been found in highly controlled laboratory environments, extend to real-world episodic memories. We ran experiments that assessed memory for experiences created in the laboratory (Experiment 1) and, using smartphones, memories for everyday experiences (Experiment 2). We complement these design innovations with a novel analysis approach to model memory accuracy and integration in both settings. Consistent with past findings, emotional events were subjectively remembered more strongly. These studies also revealed that features of more positive events were indeed more integrated within memory, both in the laboratory and the real world. These effects were specific to participants’ emotional responses to the events during encoding rather than general emotional states at the time of retrieval, and reflected a general increase in integration between multiple memory features. Together, these results demonstrate robust differences in memory for positive events, introduce a novel measure of memory integration, and highlight the importance of assessing the impact of emotion on memory beyond the laboratory.
Emotionally meaningful experiences inform our sense of self and influence future actions, with memories for such experiences diverging from the way we recall neutral events. Emotional experiences can differ in how accurately they are recalled and how integrated their myriad features are in memory. For instance, you might remember seeing a car accident on the highway. Having accurate memory for a feature of the event (e.g., where the accident occurred) could help you drive more carefully at this exit in the future. You might also have linked this feature (location) to further features of that event, like heavy rain and dense traffic, forming an integrated memory that would help you understand why this accident likely occurred and help you target your careful driving to similar conditions. Memory integration can determine how readily features of the event come to mind when cued with a reminder, and how different stimuli may cue the same event memory (Bisby et al. 2018). This phenomenon is particularly beneficial in a real-life context because more cues can bring an emotional, informative memory to mind. Critically, remembering features of an event accurately does not necessarily correspond to remembering that event holistically; these different aspects of memory involve distinct neural circuits (Davachi 2006) and may be potentiated by distinct neurotransmitters, behavioral goals (Clewett and Murty 2019), or emotional states (Bisby and Burgess 2017). Furthermore, although most research to date examining memory integration for emotional events has focused on negative experiences, there are preliminary indications that positive events may show a distinct pattern. Here, we aim to investigate the accuracy and structure of memory for emotional events, both positive and negative, inside and outside the laboratory.
A long history of work shows that positive and negative emotional experiences hold an advantage in our memories (Hamann 2001; Williams et al. 2022). People often have a stronger subjective sense of memory for emotional relative to neutral events. For example, participants were more confident that they would remember pairmates of negative compared to neutral words, even if their performance did not actually improve (Zimmerman and Kelley 2010; Caplan et al. 2019), and report a richer, more detailed sense of memory for emotional events (both positive and negative) relative to neutral (Ochsner 2000; Phelps and Sharot 2008).
On the other hand, findings are mixed regarding how accurately features of such emotional events are remembered. The central (or relevant, salient) features of negative events are often remembered more strongly than central features of neutral events (Kensinger and Corkin 2003; Kensinger et al. 2007; Bisby et al. 2018; Palombo et al. 2021), even if peripheral details (like a colored border) are forgotten (Rimmele et al. 2011). Yet the effects of centrality on memory accuracy are mixed; other studies have found such peripheral details can be enhanced as well, with reports of better memory for colored borders of emotional words (both positive and negative) relative to neutral (see Doerksen and Shimamura 2001; for negative vs. neutral only, see Kensinger and Corkin 2003). Effects can also vary based on the type of negative stimulus, with memory benefits observed for some (e.g., scenes) but not others (e.g., faces) relative to neutral (Anderson et al. 2006). Finally, negative events have been shown to be remembered with more “gist” (or less precision) compared to neutral events (Adolphs et al. 2001; Bookbinder and Brainerd 2017). Although less work has examined the effects of positive valence on memory accuracy, there is evidence that both positive and negative events are uniquely remembered compared to neutral, indicating an overall effect of emotionality (e.g., Talmi et al. 2007; for review, see Williams et al. 2022). There are, however, reports that memory accuracy differs between positive and negative events (for review, see Bowen et al. 2018), though the directionality can vary. For instance, Yegiyan and Yonelinas (2011) found that more positive (relative to more negative) stimuli enhanced memory for peripheral features. Conversely, Xie and Zhang (2017) showed that negative, but not positive, stimuli led to more precise memory for peripheral features (here, the color and orientation of a subsequent neutral object). Similarly, some studies reported increased false memory for positive compared to negative word lists (Storbeck and Clore 2005), whereas others showed the opposite (Brainerd et al. 2008). Together, the extant literature highlights complex and varying effects of valence and emotionality on the accuracy and precision of memory for different event features.
Beyond accurately remembering individual features of an event, integrating these features in memory is particularly important for adaptive behavior. As mentioned above, integration can determine how readily myriad features of the event come to mind when presented with a reminder (Bisby et al. 2018), how broadly a memory is generalized from the specific context in which it occurred (Loetscher and Goldfarb 2024), and how strongly it influences later decisions (Murty et al. 2016). There is also evidence that integrated memories may allow for more creative thinking and aid in problem solving (Isen 2015). Despite the utility of integrated memory, and the strengthening of memory for some negative event features described above, negative events have repeatedly been shown to be less integrated in memory compared to neutral events (Bisby and Burgess 2014; Bisby et al. 2018; Palombo et al. 2021). This trade-off has been used to explain problematic memories of traumatic events, which are thought to be inadequately bound to their respective contexts (Bisby and Burgess 2017), and is associated with shifts away from hippocampal-dependent integration toward amygdala-dependent potentiation of individual features (Bisby et al. 2016). More generally, distinct neural mechanisms support episodic memory for negative compared to neutral events, suggesting a pathway by which negative valence can modulate memory integration (Murray and Kensinger 2014; Yonelinas and Ritchey 2015). In contrast, emerging research suggests that positive emotions may lead to increased memory integration. In a direct test of this idea, participants could better recall cued pairmates when both were positive relative to neutral (Zimmerman and Kelley 2010; Madan et al. 2019). Recent findings corroborate these distinctions, with negative shifts in valence separating event representations and positive shifts binding them together (McClay et al. 2023), and negative, but not positive, emotion disrupting binding between event features (MacKenzie et al. 2015). Positive emotion is also associated with increased behavioral activation, a motivational state related to increased energy and exploration (Custers and Aarts 2005). Critically, behavioral activation may support greater memory integration by expanding cognitive processing and increasing the features an individual attends to, a mechanism by which positive emotion may enhance memory integration (Clewett and Murty 2019). Together, this work suggests that memory integration may meaningfully diverge between positive and negative experiences.
Finally, it is imperative to quantify the accuracy and integration of memories for real emotional experiences. As with controlled stimuli, autobiographical memory for real emotional events, both positive and negative, has been shown to be recalled with more detail and clarity relative to neutral events (St Jacques and Levine 2007; Wardell et al. 2021), with some studies finding positive events contain more peripheral details compared to negative events (Talarico et al. 2009). Other studies have examined memory for collective events with varying emotional appraisals (e.g., sports and presidential elections), but the effects of valence on these memories are mixed (Kensinger and Schacter 2006; Chiew et al. 2022). In one recent example, researchers assessing emotional states during the COVID-19 pandemic found that more negative affect predicted enhanced recall (Rouhani et al. 2023). Although these designs provide key ecologically valid insights, as the “ground truth” of such experiences cannot be known, they preclude the possibility of determining memory accuracy and often rely on retrospective determinations of affect. Furthermore, such approaches have not yet been used to assess memory integration.
The goal of the current study is to assess accuracy and integration within memories of emotional experiences, both in the laboratory (Experiment 1) and the real world (Experiment 2). In both experiments, we hypothesize that positive and negative valence will have opposite effects on memory integration, with features of more positive events becoming more integrated and features of more negative events becoming more fragmented. Given conflicting findings regarding the effects of emotional valence on accurate memory for individual features, we anticipate that emotionality (both positive and negative) may enhance memory for some features of emotional events. Finally, we hypothesize that participants will have higher subjective memory for both positive and negative compared to neutral events.
For the laboratory investigation (Experiment 1), we use a standard episodic encoding task followed by memory retrieval tests 24 h later (Fig. 1A,B). This task involves forming memories for object/scene pairs (50% of the objects portraying appetitive stimuli and 50% common household objects), and each participant reporting their emotional responses when encoding each pair. For the real-world investigation, we use a smartphone-based protocol following techniques used in ecological momentary assessment (EMA). In this design, participants are asked to report events they are experiencing in real time when they receive a notification from their smartphone, followed by retrieval tests 24 h later in which they are asked to repeat their responses (Fig. 1D,E). This approach enables us to assess affect at the time of encoding, maintain a consistent encoding/retrieval delay, and quantify memory accuracy relative to the events reported as they are occurring.
Experimental design and memory quantification. (A) In-lab memory design (Experiment 1). During encoding, participants viewed a set of 80 unique object/scene pairs. The next day, they were tested on their memory for these pairs (for exact questions, see Table 1). (B) Schematic of multifeature memory structure for Experiment 1. (C) Example of co-accuracy computation. (Left) Example distributions of a participant's accuracy scores per event feature across all trials. Squares indicate participant mean; dots indicate accuracy score on example trial. Accuracy scores, for example event n, were z-scored according to the participant's average accuracy score for the construct. (Right) We then computed the dot product per feature pair to determine element-wise co-accuracy. As shown in this example, an event could have a high co-accuracy score both when participants accurately remembered both features, and when participants forgot both. (D) Real-world memory design (Experiment 2). During encoding, participants received randomly timed prompts asking them to record the current time, their current emotional state, and aspects of their day at the current moment (for exact questions, see Table 2). The next morning, they were cued with the times from the encoding surveys and asked to remember each aspect of that event using the same questions as encoding. (E) Schematic of multifeature memory structure for Experiment 2.
Memory questions and response options for in-lab study (Experiment 1)
Questions and answer options for EMA surveys (Experiment 2)
To analyze memory integration in both experiments, we developed a novel measure of integration derived by assessing the relationship between accuracy for different event features. A key principle of memory integration is the assumption that memory for different features of the same event will be related (Horner et al. 2015). We use “integration” as a general term for coupling between event features that does not imply a particular type of question structure. This construct is often measured using targeted memory tasks, such as cueing the participant with one feature and assessing whether they can recall or recognize another feature from the same event, or having the participant choose between the original event and a version that violates the original spatial, temporal, or associative structure (for review, see Konkel and Cohen 2009). Here, we instead derive integration based on memory accuracy for different features. To provide an intuition for this measure, one can imagine correlating how well a participant remembers one feature of an event (e.g., location) with how well they remember another feature of the event (e.g., the weather) to provide a sense of holistic recollection for that individual. Critically, we can decompose these correlations across experiences into a single measure per experience. By taking the dot product of two scaled accuracy measures, we can compute the extent to which memory for one feature tracks memory for another feature for a given experience (see Faskowitz et al. 2020 for an analogous approach to fMRI data, which quantifies functional coupling per moment using edge time series; also illustrated in Fig. 1C). This “co-accuracy” technique enables us to quantify the extent to which memories for different event features are related based only on memory accuracy per feature, providing a measure of multifeature integration that can be used when assessing memory accuracy both in the laboratory (schematic of memory structure in Fig. 1B) and the real world (Fig. 1E). Using this calculation, we hypothesize that we will observe greater memory integration for more positive emotional events.
Results
Experiment 1
Accuracy
Overall, participants successfully remembered all event features (comparison to chance—item: t(42) = 7.88, P < 0.001, d = 1.2, context recognition: t(42) = 6.51, P < 0.001, d = 0.99, valence: t(42) = 21.66, P < 0.001, d = 3.3, arousal: t(42) = 12.86, P < 0.001, d = 1.96) except for context recall (recalling whether an object was paired with an indoor or outdoor scene; t(42) = 0.62, P = 0.54; Fig. 2A; for questions, see Fig. 1A and Table 1).
Experiment 1 results. (A) Memory accuracy per event feature. Gray lines = chance memory performance per feature; violin plots = actual memory performance per trial; dot = mean memory performance per feature. (B) Memory accuracy per emotion experienced in that event. Dots = mean memory performance per valence rating across participants, error bars = standard error of the mean (SEM) across participants, curves = quadratic model fits (associated with general effects of emotionality; shaded area = 95% CI). (C) Average memory co-accuracy per event feature pair. For example, the top left square shows average co-accuracy value between context recognition and valence. (D) Memory co-accuracy per emotion experienced in that event. As in B, dots show mean memory performance per valence rating, error bars = SEM, lines = linear model fits (probing effects of valence; shaded area = 95% CI). Co-accuracy calculated per feature, see (E) for all pairwise comparisons. (E) Co-accuracy per feature pair, separated by valence. (*)P < 0.05, (**)P < 0.01, (***)P < 0.001.
Co-accuracy
Using memory accuracy per feature, we computed “co-accuracy” between pairs of features within each event. The extent to which
features were integrated varied between feature pairs (main effect feature:
; visualized in Fig. 2C). Observations of these patterns provided face validity to the co-accuracy measure, as questions that probed similar constructs
also had higher co-accuracy scores. For example, we found the highest co-accuracy scores between questions probing memory
for emotional states (arousal and valence; mean co-accuracy across participants = 0.1 [SE = 0.02]) and for associated scenes
(context recall and context recognition; 0.09 [0.02]). In an exploratory analysis, we also found that co-accuracy scores were
higher for trials on which participants had correct context recognition, a standard assay of associative memory (Supplemental Fig. S2; see also Bisby and Burgess 2014, Goldfarb et al. 2019, 2020, and Sherman et al. 2023 for examples using this standard assay). Together, these results suggest that co-accuracy can provide a window into the extent
to which features of an event are integrated in memory, even without asking targeted questions presenting feature pairs.
Variability in emotion ratings
The stimuli presented in this study (randomly paired sets of objects and scenes) have previously been shown to evoke positive, negative, and neutral feelings in participants (Goldfarb et al. 2020; Sherman et al. 2023). Consistent with this, participants in the current study reported different emotions when imagining interacting with these object/scene pairs. As the objects included both neutral handheld objects (50% of pairs) and appetitive alcoholic beverages (50% of pairs), most events were rated as neutral (mean = 37.2% [SD = 14%]), followed by positive (26.7% [11%]) and negative (14.9% [9%]). Participant-level data are shown in Supplemental Figure S1A (note that one participant rated everything as “neutral” and was thus excluded from analyses probing the effects of valence).
Based on these ratings, we assess the impact of emotion at encoding on memory in two ways in each model. To assess whether rating an event as more positive is associated with altered memory for that event, we include a term for valence (with more negative values indicating more negative valence and more positive values indicating more positive valence). To test whether rating an event as inducing an emotional response (either positive or negative) is associated with altered memory for that event, we include another term for emotionality (with more positive values indicating more emotional events, both positive and negative, and 0 indicating neutral events).
Accuracy of memory for emotional events
We first tested whether accuracy differed based on emotionality and valence. Given the different scales for measuring accuracy
across features, we probed the effects of emotion on accuracy for each feature separately. For most features, we did not find
effects of either valence or emotionality on accuracy. We did find an effect of emotionality on memory for arousal (
) such that events that were rated as more emotional were remembered less accurately (Est = −0.03[SE = 0.01], P < 0.001). Although effects were not statistically significant for other features, it is worth noting that the direction of
emotionality effects was not consistent across features (e.g., context recognition was not impaired for emotional events;
Fig. 2B). These results indicate that emotionality at encoding, rather than valence, modulates memory accuracy in a feature-specific
manner.
Co-accuracy of memory for emotional events
To address our main question regarding the co-accuracy of memory for emotional events, we tested whether memory co-accuracy
differed based on valence, emotionality, and memory feature. We found a significant effect of valence on co-accuracy (
), with events that were rated as more positive being more integrated in memory (Est = 0.02[0.008], P = 0.02; Fig. 2D,E). However, unlike accuracy, there were no overall effects of emotionality, nor was there a significant interaction between
emotion and memory dimension (both P > 0.25), indicating that valence broadly influenced memory integration across features.
Subjective memory for emotional events
Based on prior work on emotional memory, we investigated whether emotion would influence the subjective feeling of remembering
(Fig. 3A). Consistent with previous findings, participants subjectively reported that memories for emotional events were more vivid
(main effect emotionality:
). There was also an effect of valence, such that the more positive an event was rated, the more vividly it was remembered
(
; effect of valence: Est = 0.03[0.007], P < 0.001). Together, these results highlight the effects of emotion at encoding on objective and subjective aspects of memory.
Emotionality enhances subjective measures of memory. (A) Experiment 1: associations between event valence and subsequent subjective feeling of memory vividness. Plot shows mean vividness per valence rating, error bars = SE. (B) Experiment 2: associations between event valence (0 = very negative, 10 = very positive) and subsequent subjective feeling of memory confidence (0 = not at all confident, 10 = very confident). Plot shows quadratic model fit, error bars = 95% CI.
Experiment 2
Completion rates
Participants completed an average of 93% of prompted encoding surveys and an average of 84% of memory surveys, yielding 874 encoding/retrieval pairs for analysis. The brief set of multiple choice and Likert questions answered in each survey is shown in Table 2. These completion rates demonstrate the practicality and feasibility of smartphone-based momentary memory sampling.
Accuracy
We quantified accuracy based on the changes in responses to the questions from the encoding compared to memory surveys (for details on how this was coded, see Materials and Methods). As in Experiment 1, participants successfully remembered event features using momentary memory sampling (comparison to chance—affect: t(42) = 46.27, P < 0.001, d = 7.06; energy: t(42) = 26.06, P < 0.001, d = 3.98; body temp: t(42) = 29.46, P < 0.001, d = 4.49; location: t(42) = 31.48, P < 0.001, d = 4.80; activity: t(42) = 22.12, P < 0.001, d = 3.37; company: t(42) = 19.66, P < 0.001, d = 3.00; Fig. 4A).
Experiment 2 results. (A) Memory accuracy per event feature. Gray lines = chance memory performance per feature; violin plots = actual memory performance per event; dot = mean memory performance per feature. (***) P < 0.001. (B) Memory accuracy per emotion experienced in that event. As in Figure 2, curves = quadratic model fits (probing effects of emotionality; shaded area = 95% CI). Dots = mean memory performance per valence ratings (for visualization only, showing average of two consecutive valence values; the model included memory per each valence rating from 0 to 10); error bars = standard error of the mean (SEM). (C) Average memory co-accuracy per event feature pair. For example, the top left square shows average co-accuracy value between location and valence. (D) Memory co-accuracy per emotion experienced in that event. As in B, dots show mean memory performance per valence ratings, error bars = SEM, lines = linear model fits (probing effects of valence; shaded area = 95% CI). Co-accuracy calculated per feature, see (E) for all pairwise comparisons. (E) Co-accuracy per feature pair, separated by valence. For the purpose of visualization, valence ratings of 0–3 are coded as negative, 4–6 as neutral, and 7–10 as positive. (*) P < 0.05; (**) P < 0.01; (***) P < 0.001.
Co-accuracy
We computed co-accuracy between feature pairs (see Fig. 4C) and found a main effect of memory feature (
). We found the greatest integration between location and activity (0.18 [0.04]) and the lowest integration between company
and energy (−0.06 [0.04]).
Variability in emotion ratings
All participants showed a range of valence ratings at encoding (where 0 = very negative and 10 = very positive; ratings per participant ranged from an average minimum of 2.58 [SD = 1.52] to a maximum of 8.33 [1.74]). Variance per participant is shown in Supplemental Figure S1B. We note that events that were rated as more positive, while also associated with higher energy (Est = 0.53[0.04], P < 0.001), temperature (Est = 0.21[0.06], P < 0.001), and particular locations (F(10,167.43) = 2.08, P = 0.03; most positive affect at a hotel and least positive affect at work), did not systematically differ in other features that would likely influence the way they were remembered (e.g., number of companions, activities, time of day, whether this was a relatively usual or unusual event).
Accuracy of memory for emotional events
We tested whether memory accuracy for different features of real-world events differed based on the emotionality of the event.
As in Experiment 1, we did not find consistent effects of valence or emotionality at encoding on memory accuracy across features.
We again found an effect of emotionality on accuracy for one memory feature, temperature (
; see Fig. 4B), with worse accuracy for more emotional events (Est = −0.002[0.0006], P = 0.02). These results show that emotion had differing impacts on memory accuracy depending on the memory feature, even in
a real-world setting.
Co-accuracy of memory for emotional events
As with accuracy, we assessed whether memory co-accuracy was impacted by emotionality and valence. Critically, we found a
main effect of valence on memory co-accuracy (
), such that the more positive an event was rated, the more it was integrated in memory (Est = 0.02[0.005], P < 0.001; Fig. 4D,E). This valence-induced increase in co-accuracy was not influenced by feature (valence × feature: P > 0.25). There was no significant effect of emotionality (main effect and interaction with features; both P > 0.25). Together, these results indicate selective benefits in memory integration for more positive events—and conversely,
less integration for more negative events—in the real world.
Subjective memory for emotional events
As participants reported their confidence in their memories, we also investigated whether emotion influenced the subjective
feeling of remembering (Fig. 3B). We found a significant effect of emotionality on confidence (
, such that more emotional events (both positive and negative) were remembered more confidently (Est = 0.004[0.001], P = 0.003). There was no significant effect of valence (P > 0.25), indicating that both positive and negative events were remembered more confidently than neutral.
Influence of emotional state at retrieval
In addition to the emotional content of the memory, it is possible that an individual's emotional state at retrieval would influence how they remember past events (for review, see Kensinger and Ford 2020). Prior to completing their memory surveys or being exposed to any memory cues, participants reported how they were currently feeling (very negative to very positive), enabling us to assess how general affective states immediately prior to retrieval were related to memory. As with the analyses above, we included separate terms for valence (assessing whether feeling more positive prior to retrieval was associated with differences in memory) and emotionality (assessing whether generally feeling more emotional prior to retrieval was associated with differences in memory) as potential predictors of memory.
There were no significant effects of valence or emotionality at retrieval on accuracy for any feature individually. We next
tested whether the emotional state at retrieval impacted co-accuracy. Notably, unlike valence at encoding, we did not find
a significant effect of valence at retrieval (F(1,745.67) = 2.86, P = 0.09; interaction with memory feature P > 0.25). We did find an overall effect of emotionality (
), with higher co-accuracy if feeling more emotional at retrieval (Est = 0.09[0.05], P = 0.04; interaction with memory features P > 0.25).
Together, these results demonstrate that emotional states at both encoding and retrieval can influence real-world memories, consistent with extant literature (Kensinger and Ford 2020), and that valence at encoding is particularly important for determining co-accuracy of these events.
Discussion
Here we investigated the structure of memories for emotional experiences in the laboratory and the real world. Across two experiments, we found that more positive valence at encoding was associated with more integrated memories the next day. This work provides new experimental and analytic tools to quantify memory organization and highlights the ways in which emotional experiences are remembered differently.
Our finding that more positive valence at encoding increased memory integration spanned both in-lab and real-world settings and was consistent across all memory features measured. These results are consistent with recent literature showing that positive emotion is associated with increased integration (Zimmerman and Kelley 2010; Madan et al. 2019; McClay et al. 2023; Sherman et al. 2023). They also align with models indicating that positive valence leads to broadened thinking (Fredrickson 2004), triggering relational processing and promoting “integration of different types of information in the environment” (Kaplan et al. 2012). Conversely, particularly in a real-world setting, we found that more negative emotion disrupted memory integration, again consistent with past laboratory findings (Bisby and Burgess 2014, 2017; Bisby et al. 2018). These findings provide a key extension to past work by showing that greater memory integration for positive events was evident across settings, memory assessments, and stimuli. This consistent pattern underscores the need for further research to explicate the neural and cognitive mechanisms by which positive valence drives integration (for recent proposals, see Bowen et al. 2018; Clewett and Murty 2019), and to probe how valence may influence integration between as well as within episodes (Loetscher and Goldfarb 2024). We note that, as the participants themselves determined their own emotional responses to stimuli (Experiment 1) and events (Experiment 2), we could not fully match nonaffective features of the memoranda that induced positive or negative emotional responses. Indeed, in standard studies of autobiographical memory, the events that participants describe as positive and negative differ in many features in addition to valence (e.g., common positive events in one study included parties, whereas negative events included arguments with relatives; see D'Argembeau et al. 2003) and the laboratory stimuli that evoke different emotional responses can also vary (Vuilleumier et al. 2003; Wilms and Oberfeld 2018). We verified that valence ratings in Experiment 1 were not significantly related to semantic structure or normed ratings of detail and familiarity, and that valence ratings in Experiment 2 were not significantly related to novelty or time of day (see Talmi et al. 2007; Poppenk et al. 2010 for discussion). Furthermore, the consistency of the effects of valence on memory co-accuracy, despite many differences in the types of events in the laboratory versus the real world that evoked these emotional responses, suggests a relationship between valence and memory integration that extends beyond the features of the events themselves. Nevertheless, more work is needed to determine how the phenotypes of experiences that evoke different emotional responses are related to differences in how these experiences are remembered.
In addition to demonstrating distinct memory organization for positive emotional events, this work presents a new behavioral measure of memory integration. This “co-accuracy” measure can be broadly applied across any data set that includes accuracy values for ≥2 features of an event, in contrast to extant measures that require specialized memory assessments and reexposure to select event features. In Experiment 1, we showed that event co-accuracy was related to a standard laboratory assay of associative memory. In Experiment 2, we showed that co-accuracy could be applied to a real-world setting and to an event with five features, for which integration would be challenging to assess using standard pairwise cue presentations. Co-accuracy thus provides a scalable and effective tool that captures the extent to which a feature is bound to others in memory. Interestingly, we also found that location was the most integrated memory feature. This is consistent with the primacy of spatial information in models of hippocampal-dependent episodic memory (Burgess 2002) and event binding (Davachi 2006).
By applying this same analysis technique to both laboratory-based and smartphone-based data, we could begin to examine general patterns in how features of episodic memories are integrated across settings. This approach provided robust and ecologically valid evidence that more positive events differ in their memory structure. It also provides intriguing preliminary evidence that memory structure may differ across settings, as we found overall higher co-accuracy for real-world (Experiment 2) compared to laboratory-based (Experiment 1) memories. It is important to note that this seeming increase in integration may be related to the distinct memory features assessed in the smartphone-based (e.g., activities, companions, locations) compared to the laboratory-based (e.g., items, contexts) experiment. These real-world events likely co-occur more often than simulated laboratory events and inherently have more personal relevance and meaning. On the other hand, these considerations may reflect a meaningful difference in the quality of memories that we can capture in the laboratory versus the real world. Autobiographical memories have been shown to involve greater self-referential processing (Cabeza et al. 2004; Cabeza and St Jacques 2007), which can in turn promote memory binding or integration (Sui and Humphreys 2015). Furthermore, autobiographical memories are richer in sensory details (Cabeza and St Jacques 2007), which may allow for greater integration between these different details. Although more research is needed to test these possibilities, this analysis approach is a practical and scalable way to analyze both real-world and in-lab memories and begin to assess potential differences in memory structure between these settings.
The current set of experiments replicates previous findings regarding subjective memory for emotional events. For example, we found in both studies that more emotional events were remembered more vividly (Experiment 1) and confidently (Experiment 2) (Talarico and Rubin 2003; Kensinger et al. 2011; Rimmele et al. 2011). As in past reports, we found that subjective memory was broadly increased by emotionality rather than specifically enhanced by positive or negative valence (although we note that, in Experiment 1, there was a particular advantage for vividly remembering positive events). Notably, and again consistent with past research, this subjective sense of remembering was not equivalent to more accurate memory for event features. Instead, across studies, some features of emotional events were remembered less precisely (consistent with past reports of more “gisty” memory for emotional events; see Adolphs et al. 2001; Bookbinder and Brainerd 2017). In contrast to past reports, we did not see significant enhancements in accuracy for features of emotional events. It is possible that our stimuli were not sufficiently emotional: a margarita will not evoke as strong a reaction as an image of an injured person (Talmi 2013; Bisby et al. 2016), and randomly prompting someone during their day is unlikely to capture experiences as intense as a global pandemic (Rouhani et al. 2023). Nevertheless, the emotion induced here was sufficient to alter the integration of features within a memory. This is consistent with past reports showing that negative affect can impair associative memory even without significantly modulating memory for individual features, which may reflect a greater sensitivity of integration to emotional valence, although further work is needed to test this idea in the context of positive valence (Bisby and Burgess 2014; Bisby et al. 2018).
More broadly, a major challenge in probing the effects of valence on memory is the concomitant influence of arousal, or the degree to which an emotion is exciting/calming (Russell 1980; Scherer 2005). Arousal can influence the allocation of attention at encoding (Easterbrook 1959; Kensinger and Corkin 2003) and prioritize distinct memory features (Anderson and Phelps 2001; Mather 2007; Clewett and Murty 2019). Positive and negative experiences can also differ in arousal, with laboratory experiments often showing negative images that are more arousing; thus, there is ongoing debate about how much valence or arousal contributes to alterations in memory (Mather and Sutherland 2009; Bowen et al. 2018; Clewett and Murty 2019). In our studies as well, positive and negative events were not matched in intensity, although the direction is atypical. Here, rather than negative stimuli being more arousing, positive events in Experiment 1 were rated as more emotionally intense than the negative events, and valence ratings in Experiment 2 were skewed such that positive events were rated as more strongly positive than the negative events were negative. Notably, these (more arousing) positive events were the same ones for which we observed greater memory integration. This finding diverges from models positing that more arousal would lead to more fragmentation (Clewett and Murty 2019). Although our positive events may not have been sufficiently arousing to disrupt integration, this finding may suggest an interactive effect of positive valence and high arousal in promoting the integration of features within events (see also Steinmetz et al. 2010).
The current study provides a compelling argument for the introduction of smartphone-based memory techniques, similar to those used for EMA in clinical settings, to emotion and memory research. This momentary memory assessment approach presents several advantages to extant measures of autobiographical memory, as we can assess whether memories are recalled accurately by comparing retrieval to initial encoding reports, measure event characteristics (e.g., affective state) at the time of encoding, and control for the delay between encoding and retrieval. We also note that momentary sampling, as described here, addresses some challenges encountered in daily diary studies used to assess memory. First, these designs frequently require participants to report events at the end of each day (Conner and Silvia 2015; Schnitzspahn et al. 2016; Anderson and Fowers 2020; Brown et al. 2021). This requires the participant to engage in a reflective memory search during this initial event reporting, a retrospective process that is subject to potential biases. Second, there are challenges in diary studies regarding how to cue participants as to which memory you want them to retrieve. Possible approaches include using photos (Pathman et al. 2013) or geographical location (MacKenzie et al. 2020) as part of a memory test, or potentially encouraging participants to assign a title or keyword to an event as they report it. However, such cues may themselves contain information and trigger associations, thus augmenting memory retrieval (and, with particularly detailed cues, even altering hippocampal representations; see Martin et al. 2022). Although this limitation can be addressed using free recall without specific cues, coding and interpreting such responses is complex, and it may not be accurate to attribute information that was not reported to being forgotten. Using timestamps as the memory cue in the current study and explicit prompts regarding each memory feature allows us to direct the participant to the corresponding event while avoiding the presentation of event details.
We also note that there are limitations to the current momentary memory assessment techniques. The act of reflecting on an event inherent to recording it could encourage elaboration of the event, which aids in connecting details of the event with prior knowledge and enhances memory (Kensinger 2004). Even with the relatively minimal reflection required for the current study (responding to a short set of Likert and multiple choice questions), it is possible that the experience of recording would alter the encoding process (Cabeza and St Jacques 2007). Additionally, the act of responding to smartphone prompts may become a feature of the event, meaning that characteristics at encoding (e.g., affect, stress) could be influenced by the experience of this sampling method, which could in turn influence memory (MacKenzie et al. 2020). Such limitations could be addressed by incorporating passive monitoring (such as GPS tracking or physiological measures such as heart rate; see MacKenzie et al. 2020). In addition, our question format, while minimizing participant burden and allowing for computation of accuracy, precludes analyses possible with free text responses, like examining the sentiment or type of detail included in a memory (Rouhani et al. 2023). Future studies could augment the prompts presented here with free recall to capture these patterns. Finally, we are limited to the events that occurred when we sent our prompt. Other events (potentially including ones with greater emotional salience) could not be captured. This limitation too could be addressed by incorporating ongoing passive monitoring using sensors (MacKenzie et al. 2020). Further work is needed to determine the optimal number and timing of encoding prompts to ensure high completion rates and potentially enhance the recording of emotional experiences.
In summary, we demonstrate that features of more positive events are more integrated in memory. We present a new behavioral measure of memory integration, deemed “co-accuracy,” which is widely applicable across data sets, and we demonstrate its construct validity. We replicate our findings in a highly controlled laboratory environment and a real-world setting utilizing smartphone data collection technology, showing the practicality of this approach and illustrating the value it adds to the emotional memory field. Emotional memories are core to our sense of self, and influence our goals and how we move through the world. Our work contributes to an understanding of how positive and negative feelings transform the way that these experiences live on in our memories.
Materials and Methods
Participants
Fifty healthy adult participants aged 18–45 years old were recruited from the New Haven community and provided consent to participate in both experiments. We excluded data from four participants who did not complete procedures from both experiments (N = 2 did not complete retrieval for Experiment 1; two did not complete Experiment 2) or provided insufficient smartphone-based data for analysis (≤2 responses; N = 3). We therefore had a final sample size of 43 participants (age: mean = 24.4 [SD = 5.59], sex assigned at birth: 53% female, 47% male).
Both Experiments 1 and 2 included a 1-day delay between encoding and retrieval, allowing for sleep and consolidation of these long-term episodic memories and standardizing the delay between experiments (for evidence that delay can modulate effects of valence on memory integration, see Pierce and Kensinger 2011).
Experiment 1
Experimental design and procedure
Experiment 1 took place over 2 days in the laboratory, with encoding on day 1 and retrieval on day 2 (for similar procedures, see Goldfarb et al. 2020; Sherman et al. 2023). All procedures for both experiments were approved by the Yale University IRB.
Stimuli
Stimuli were described in detail in Sherman et al. (2023). In brief, photographs of alcoholic beverages and handheld objects were obtained from prior studies (Dunsmoor et al. 2012; Van Der Linden et al. 2015; Fey et al. 2017; Sinha et al. 2022) and from Google image searches. Each object was edited to appear in a gray square (250 × 250 pixels) with text occluded and was selected to be maximally distinct from other category members. Individual objects were validated in a separate online experiment as inducing varying levels of valence and arousal (Bradley and Lang 1994), and were rated on levels of perceptual detail (Dager et al. 2014) and familiarity (Bainbridge et al. 2017). For each participant, a subset of this 400-image corpus was selected such that alcoholic beverages and handheld objects were matched in familiarity and detail, and that images presented in item recognition tests (old vs. foils) were matched in valence and arousal.
For scenes, indoor (N = 40) and outdoor (N = 40) photographs were obtained from the SUN database (Xiao et al. 2010) and Google image searches. Each scene had a same scene-type perceptual match that was presented during the context recognition task (see Fig. 1A). A separate cohort rated the similarity of each scene/match pair. For each participant, a subset of the full 320-scene corpus was selected and randomly paired with objects such that perceptual similarity was equated between the alcohol and handheld object blocks and each block contained an equivalent number of indoor and outdoor scenes.
Encoding
On day 1, participants encoded 80 trials, each featuring a randomly generated pairing of a photograph of a scene together with a photograph of an object on a gray background. These trials were presented over two blocks. In one block, the objects portrayed alcoholic beverages; in the other block, the objects portrayed handheld household objects (order counterbalanced). The two categories of stimuli had a similar underlying semantic structure, with each divided into four subcategories (i.e., alcoholic beverages were equally subdivided into wine/beer/liquor/mixed drinks; handheld household objects were equally subdivided into items found in the kitchen/bathroom/garage/office).
On each trial, participants were instructed to imagine interacting with the object in the scene as vividly as possible (5 sec). Participants were aware there would be a memory test the next day, and were advised that people tend to do better on the memory test when they imagine the interaction more vividly. Participants were then asked to rate how they were feeling when imagining that object/scene pair, first providing their valence (happy, unhappy, neutral; 2 sec), how intensely they were feeling that way (not at all to very; 2 sec), and how much they wanted an alcoholic beverage (not analyzed here; 2 sec).
We note that valence was not predetermined by the experimenters. Rather, valence was assigned per trial based on participants’ ratings of how they were feeling when viewing that specific object/scene pair. In past studies, these object/scene pairs have been shown to elicit high arousal states and to be interpreted as positive or negative (Goldfarb et al. 2020; Sherman et al. 2023). We confirmed that objects assigned different valence ratings did not differ significantly based on normed ratings of levels of perceptual detail and familiarity (detail: P = 0.15; familiarity: P = 0.96), or luminance (P = 0.88). We also did not find a significant relationship between valence ratings and object subcategories (relationship between valence and proportion within subcategory: P = 0.58). Follow-up analyses confirmed that object category did not significantly predict co-accuracy or accuracy.
Retrieval
The next day, participants returned to the laboratory and responded to questions regarding their memory for these object/scene pairs. First, they were shown individual images of objects (160 objects total, 50% novel, viewed for 3 sec each) and asked to respond to whether these objects were seen during encoding (old) or not (new, 2 sec response window). This served as our measure of item recognition. Next, participants were shown objects from encoding (2 sec each) and asked to select whether they had been paired with an “indoor” or “outdoor” scene (during encoding, 50% of paired scenes were indoor and 50% outdoor; 2 sec). We label this as context (recall). Then, participants were shown the objects from encoding along with a set of four scene pictures. These scene pictures included: the correct pairmate from encoding, a perceptually similar lure, another encoded scene (but not the correct pairmate for that object), and its perceptually similar lure (4 sec to respond). We label the selection of the correct pairmate from encoding as context (recognition). Finally, we assessed memory for affect at encoding. Participants were shown each object/scene pair from encoding (2 sec). They were then asked to remember how they felt when imagining that object/scene pair the prior day, and to repeat their ratings of valence, arousal, and craving based on how they remembered feeling (2 sec/response, as at encoding). Differences in ratings from encoding to retrieval provided measures of memory for valence and arousal (for more details on memory computation, see “Defining memory accuracy” below). As a subjective measure, participants rated how vividly they remembered their feelings from encoding (not at all vivid to extremely vivid; 2 sec per response).
Experiment 2
Participants
The subject pool was identical between Experiment 1 and Experiment 2 to facilitate comparisons. We note that no participants were excluded based on not having reliable access to a smartphone or inability to download the MetricWire app.
Experimental design and procedure
Smartphone application
The current study used the smartphone survey application MetricWire (app versions 4.9.8–4.10.3). Participants were trained during an in-person session on how to use the application. After downloading the application onto their smartphones, participants completed a “Set Up Survey” with information such as the time they normally wake up to appropriately time survey prompts.
Prompt structure
Participants received two prompts on their smartphones every day (hereafter referred to as encoding surveys) during the 2-week monitoring period. One encoding survey was sent in the morning (randomly within 3 h after wake-up time as recorded in the setup survey) and one later in the afternoon (randomly between 3:00 p.m. and 7:00 p.m.). After receiving the notification, participants had a window of 90 min to complete each survey. Participants received another notification after 15, 30, and 60 min if they did not complete the survey. Each encoding survey began with asking participants to enter the current time and to describe what is happening “right now” in response to a brief set of multiple choice and Likert questions (see Table 2 for all questions presented).
Participants were prompted at their reported wake-up time the following morning to complete up to two memory retrieval surveys, triggered based on successful completion of the previous day's encoding surveys. As with encoding surveys, participants had 90 min to complete these surveys, and received notifications after 15, 30, and 60 min if not completed. The memory surveys pipe in the time entered by participants during the associated encoding survey from the previous day to serve as a memory cue. For example, a participant may be asked “yesterday at [1:00 p.m.]…” Memory surveys were identical to encoding surveys, except describing events from “yesterday” rather than “right now.” In addition, as a subjective measure of how they felt about these memories, participants reported their confidence in their ability to remember the event on a scale of 0 to 10.
Prior to completing retrieval surveys, participants completed a brief morning check in. In this survey, they were asked, “How are you feeling about today?” on a 0- to 10-point Likert scale from “very negative” to “very positive.” These responses were used to derive participants’ emotional state at the time of retrieval.
The question structure and response options were informed by past EMA studies (Carney et al. 2006; Epstein and Preston 2010; Dunton et al. 2015; Burke and Naylor 2020). We further validated and refined questions in our pilot studies to ensure that Likert questions evoked a range of responses and that options listed in multiple-choice questions were representative of the experiences of participants in our sample (e.g., listed activities that they reported engaging in while minimizing responses of “other”).
Statistical analysis
Statistical analyses were conducted in R 4.4.1. To facilitate comparison between experiments, we took a similar modeling approach
throughout. We used linear mixed effects models (lme4 package, Bates et al. 2015). F-values were calculated using type III analysis of variance with Satterthwaite's method implemented in lmerTest, and coefficient
estimates were extracted using the summary() function. Effect sizes were calculated using sjStats (
) and lsr (Cohen's d). For each model, we included subject as a random effect (due to convergence issues, we did not include random slopes). When
the main effects from these models were significant, we ran follow-up analyses without interaction terms and confirmed that
the main effects remained significant. Results were considered significant if P < 0.05. Data and analysis code are available on OSF (https://osf.io/53ygd/).
For all models, we aggregated memory performance per participant and any predictors (e.g., emotion at encoding). A separate set of analyses in which data was modeled at a trial level largely yielded the same pattern of results.
In each model, we included two continuous emotion-related predictors. The first enabled us to assess the effects of valence, specifically whether feeling more positive was associated with differences in memory. In Experiment 1, ratings were recoded as ranging from −1 to 1 (such that −1 = negative, 0 = neutral, 1 = positive). In Experiment 2, ratings were recoded as ranging from −5 to 5 (converted from 0 to 10 to maximize comparability with Experiment 1). This term thus allows us to separate the effects of feeling more positive and feeling more negative. The second term enabled us to assess the overall effects of emotionality, specifically whether having any emotional response was associated with differences in memory. This was calculated as valence2. Thus, more positive values reflect greater emotional responses (both positive and negative).
Defining memory accuracy
All memory features were scored for accuracy on a scale of 0–1, with 1 being completely accurate. Scoring differed based on experiment and memory feature as described below. Accuracy was modeled separately per feature, with accuracy predicted by emotional valence and emotionality with subject as a random effect.
Experiment 1
For item and context memory, accuracy was coded as 1 for correct responses (recognizing an object as “old,” accurately selecting whether the object was paired with an “indoor” or “outdoor” scene, and selecting the correct scene pairmate from a list of four scene images) and 0 for incorrect responses. For valence and arousal, accuracy was computed as the difference between responses at encoding and retrieval, with 1 − abs(Δ(retrieval − encoding))/numoptions, where numoptions = the length of the response scale.
Experiment 2
Broadly speaking, memory accuracy was computed as the difference between responses on the encoding and retrieval surveys for each encoding/retrieval pair. The exact calculation differed depending on the structure of the question.
For Likert scale responses (affect, energy, body state), accuracy was computed as 1 − abs(Δ(retrieval − encoding))/numoptions (where numoptions = the length of the Likert scale, or 11).
For multiple choice questions in which only one response could be selected (location), accuracy was computed in a binary fashion (1 if retrieval==encoding, otherwise 0).
For check-all-that-apply options (company, activities), we used a hypergeometric cumulative distribution function to compute whether the overlap between responses at encoding and retrieval was above chance. This calculation estimates the probability of getting the observed amount of overlap as a function of the number of overlapping responses (responses in both the encoding and retrieval pools), the number of total options available to choose from, and the number of options selected at encoding and retrieval. To make this measure comparable to our other indices (where 1 = highest accuracy), we converted this to 1-CDF. Thus, a value >0.95 indicates greater overlap than would be expected by chance, or successful memory.
Defining memory co-accuracy
Co-accuracy calculates the extent to which memory for one feature tracks memory for another feature of the event and was computed similarly across experiments. After z-scoring accuracy per memory feature, we calculated the element-wise product between pairs of memory features. This provides an estimate of co-accuracy per feature pair per memory, and is comparable to methods used in fMRI analysis to calculate edge time series (or frame-wise functional connectivity; see Faskowitz et al. 2020). This approach has the property that taking the average co-accuracy of features i,j across all events is equivalent to correlating the accuracy of memory for i with accuracy of memory for j across events. It also allows us to test the extent to which memory for one feature is dependent on memory for another feature, both by providing higher values when both features are remembered and when both features are forgotten (for another framework that similarly incorporates both co-remembered and co-forgotten feature pairs into a measure of dependency, see Horner and Burgess 2013). We note that, prior to computing co-accuracy, fully forgotten events (for which memory for all features was below chance) were excluded. Any co-accuracy values that were more than 5 SD outside the group mean for that feature pair were removed prior to analyses. For models of co-accuracy, we first computed the average co-accuracy between each event feature and all other event features (resulting in a single co-accuracy score per feature, analogous to the single accuracy score per feature described above). Co-accuracy was included in a single model as a function of feature, emotional valence, and emotionality, with subject as a random effect.
Modulation by emotion
Experiment 1
As described above, participants rated how they were feeling when encoding each trial-unique object/scene pair. For valence, they were asked whether they felt “happy,” “neutral,” or “unhappy.” We then used these ratings from encoding to create emotional valence (1,0,−1) and emotionality (1,0,1) predictors of later memory accuracy and co-accuracy using the linear mixed effects model described above. We note that, to avoid circular logic, any analyses probing the effects of valence on memory did not include memory for valence itself.
Experiment 2
During each encoding survey, participants reported how they were feeling on a Likert scale from 0 (very negative) to 10 (very positive). These ratings were then recoded by emotional valence (−5 to +5) and emotionality (0 to 5, taking the absolute value of the deviation from neutral, or 5), which were then entered as continuous predictors of later memory accuracy and co-accuracy. As above, any analyses probing the effects of valence on memory did not include memory for valence itself. To confirm that changes in memory were associated with emotional responses to the events as they were occurring, rather than participants’ general emotional states when retrieving these memories, we repeated these analyses using self-reported emotional state at retrieval to derive predictors of memory. Note that participants provided these ratings of emotion each morning immediately prior to completing up to two memory tests.
Subjective evaluation of memory
In both experiments, we assessed participants’ subjective appraisals of their memory for each event. We used the same modeling approach described above, predicting the subjective feeling of remembering as a function of emotional valence and emotionality.
Acknowledgments
The authors are grateful to Sanghoon Kang, Grace Larrabee, and DT Nguyen for assistance with data collection, and Jean Ye and Nia Fogelman for helpful discussions. This work was supported by K01 AA027832.
Footnotes
-
[Supplemental material is available for this article.]
-
Article is online at http://www.learnmem.org/cgi/doi/10.1101/lm.053971.124.
- Received August 28, 2024.
- Accepted December 22, 2024.
This article is distributed exclusively by Cold Spring Harbor Laboratory Press for the first 12 months after the full-issue publication date (see http://learnmem.cshlp.org/site/misc/terms.xhtml). After 12 months, it is available under a Creative Commons License (Attribution-NonCommercial 4.0 International), as described at http://creativecommons.org/licenses/by-nc/4.0/.














