Sleep

Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER.

TL;DR

Pooled qEEG emotion-recognition scores can reflect participant-specific recording structure rather than transferable affective information, as participant-independent evaluation of the same features and model returned near-chance performance compared to inflated epoch-pooled results.

Key Findings

Epoch-pooled evaluation on DEAP yielded substantially inflated performance compared to participant-independent evaluation of the same qEEG features and model.

  • Epoch-pooled evaluation on DEAP gave ROC-AUC values of 0.689 for valence and 0.711 for arousal.
  • Participant-independent evaluation of the same features and model returned 0.493 for valence and 0.447 for arousal.
  • Grouping epochs by trial accounted for approximately 0.06 of the performance difference, and separating participants accounted for a further 0.13 to 0.17.
  • The evaluations used the same features and model, isolating evaluation protocol as the variable driving the performance gap.

The qEEG features used for emotion recognition could identify individual participants with near-perfect accuracy, indicating they encode participant identity rather than purely affective content.

  • The same features identified participants with accuracy of 0.998 on DEAP and 0.891 on DREAMER.
  • A predictor using no EEG data—assigning each trial its participant's training-set positive rate—accounted for 42 to 84 percent of the above-chance discrimination achieved by the epoch-pooled model.
  • This demonstrates that participant-specific recording structure, rather than affective information, drives much of the apparent emotion-recognition performance in epoch-pooled protocols.

Within-participant emotion-related effects were reproducible on DEAP but close to zero on DREAMER, and effect directions were inconsistent across participants.

  • Emotion-related effects were reproducible within participants on DEAP but close to zero on DREAMER.
  • The direction of emotion-related effects reversed for approximately 40 percent of features across participants.
  • This high inter-individual variability undermines population-level claims about qEEG emotion biomarkers.

Personalized (within-participant) training provided a statistically reliable improvement over participant-independent evaluation only for DEAP valence, and not for DEAP arousal or either DREAMER target.

  • In a matched participant-level comparison using a single fixed estimator, training on a participant's own data improved DEAP valence by 0.092 AUC (95% CI 0.029 to 0.157, Holm-adjusted p = 0.042).
  • No reliable benefit was found for DEAP arousal in personalized training.
  • Neither valence nor arousal showed reliable personalization benefit on DREAMER.
  • The authors concluded that personalization should only be considered where stable within-person effects are demonstrated.

Cross-dataset replication of per-feature arousal effect sizes was modest, and no individual feature reached false-discovery-rate significance in both datasets.

  • Across channels shared by DEAP and DREAMER, per-feature arousal effect sizes correlated moderately between datasets.
  • No individual feature reached false-discovery-rate significance in both datasets simultaneously.
  • This limits the generalizability of any specific qEEG feature as a cross-dataset emotion biomarker.

Common evaluation protocols in qEEG emotion recognition allow overlapping epochs and recordings from the same participants to appear in both training and test sets, constituting a methodological confound.

  • The study identifies that overlapping epochs from the same participants in both training and test sets inflate apparent model performance.
  • Four evaluation protocols were compared: epoch-pooled, trial-grouped, participant-independent, within-participant, and cross-dataset.
  • The paper argues that population-level claims require participant-independent evaluation to avoid confounding by subject identity.

What This Means

This research examined a common problem in studies that use brainwave (EEG) measurements to recognize human emotions. Many such studies report promising results—suggesting EEG can reliably detect whether someone feels happy or excited—but this paper found that those results are largely an artifact of how the studies are designed and evaluated. When the same person's brainwave data appears in both the training and testing portions of the experiment, the computer model learns to recognize the individual person rather than their emotional state. The researchers showed this directly: the same EEG features used for emotion detection could identify individual participants with 99.8% accuracy, and a model that simply used each participant's personal baseline—without looking at any EEG at all—could explain between 42% and 84% of the apparent emotion-detection performance. When the researchers applied a stricter test—keeping data from each participant entirely out of the training set—the emotion-recognition performance dropped to essentially chance levels (around 0.49–0.45 on a 0–1 scale where 0.5 is random guessing). Emotion-related brain signals also varied enormously between people: the direction of the effect reversed for about 40% of EEG features across different participants, meaning what looks like a 'high arousal' signal in one person might look like a 'low arousal' signal in another. Only on one dataset (DEAP) and one emotion dimension (valence, or pleasantness) did personalized models—trained on each individual's own data—show a statistically reliable improvement. This research suggests that many published claims about EEG-based emotion recognition may be overoptimistic because they don't adequately separate data from individual participants during testing. The findings imply that any genuinely useful EEG emotion tool would need to be tailored to each individual person, and only after demonstrating that stable, consistent emotional signals exist within that person. Studies making broad population-level claims should use evaluation methods that test performance on entirely new, previously unseen individuals.

Check Your Own Numbers

Upload your bloodwork. We'll cross-reference your results against this study and 4,700 others.

Upload Your Labs

Have a question about this study?

Citation

Pandilova E, Stojmenski A, Chorbev I, Petrov M, Kitanovski I, Trajanov D. (2026). Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER.. Sensors (Basel, Switzerland). https://doi.org/10.3390/s26175327