MOV-AAD Dataset

A Large-Scale Multimodal Dataset for Auditory Attention During Moving Conversations

Overview

Motivation

Auditory attention decoding (AAD) is often evaluated on static, simplified speech scenes that poorly match everyday listening conditions. MOV-AAD addresses this gap by providing a large-scale dataset designed to study auditory attention under moving, naturalistic conversations. The dataset combines 64-channel EEG with synchronized physiological recordings, enabling analysis of cross-modal neural and physiological markers of attention and listening effort in ecologically valid listening scenarios.

Experimental methodology and multi-modal analysis pipeline for moving speakers
Figure 1: Experimental methodology and multi-modal analysis pipeline. The diagram illustrates the dynamic spatial movement of conversational speech sources (±90° azimuth), the synchronization of 64-channel EEG with peripheral physiological modalities, the full experimental pipeline, and the stimuli description.

Key Features

  • 64-channel EEG recordings from 50 participants
  • Synchronized multimodal physiological signals (eye tracking, respiration, GSR, PPG, SpO₂, temperature, motion)
  • Dynamic spatial movement of conversational speech sources (±90° azimuth)
  • Naturalistic conversational speech stimuli with behavioral attention tracking
  • Over 4,800 trials with approximately 45 minutes of multi-conversation tasks recordings plus 30 minutes of single-conversation tasks recordings per participant

Applications

This dataset can be used for:

  • Robust auditory attention decoding in realistic spatial dynamics
  • Multimodal attention modeling integrating neural and autonomic measures
  • Listening effort assessment using peripheral physiological markers
  • Intersubject neural response analysis during naturalistic speech
  • Benchmarking selective auditory attention algorithms

Dataset Description

Overview Statistics

9
Synchronized Modalities

64-ch EEG, Gaze location, Pupil dilation, Respiration, GSR, PPG, SpO₂, Temperature, Motion

1200 Hz
Unified Sampling Rate

Temporal alignment across all neural & physiological streams

50
Participants

Age 24 ± 4.5 years, verified normal hearing status

~ 75 min
Duration per Subject

Massive continuous recording for robust AAD modeling

4
Experimental Tasks

Repeated sentences, localization, single & multi-speaker conversations

Rich
Behavioral Tracking

Active repeated-word detection & precise localization reports

Data Composition

The dataset encompasses four distinct experimental tasks per participant, transition from structured sensory validation to ecologically valid, naturalistic listening scenarios:

1. Repeated Sentence Task: A baseline neural reliability check featuring shuffled short sentences repeated without background noise to evaluate within-subject response consistency.

2. Localization Task: A spatial perception task conducted prior to the main experiment to familiarize participants with spatial cues and assess their localization accuracy.

3. Single-Conversation Task: An active listening task where participants track naturalistic, continuous conversational narratives from a single dynamically moving sound source.

4. Multi-Conversation Task: A selective attention task requiring participants to attend to one target conversational stream while ignoring a competing co-spatial distractor.

Task / Condition Experimental Purpose Stimuli & Speakers Spatial Configuration Background Noise Total Stimuli / Duration Available Modalities
Repeated Sentence Task Split-half & odd-even reliability check for neural responses. 6 short sentences (shuffled order); 3 male & 3 female talkers. Diotic presentation (None) None 120 trials
(~12 min)
✓ EEG & Physio
✓ Behavioral
Localization Task Familiarize spatial cues & assess baseline spatial perception. Independent short sentences; Single talker per trial. Static; 9 coordinates (-90° to +90° in 22.5° steps). None 54 sentences
(Short segments)
- No Neural
✓ Behavior Metrics
Single-Conversation Task Track neural tracking of speech with dynamic spatial change. Continuous dialogue context; Multi-turn talkers. HRTF-based dynamic moving (-90° to +90°); RMS-matched. Diotic pedestrian/babble noise (-9, -12 dB) 40 trials
(~30 min)
✓ EEG & Physio
✓ Behavior & Audio
Multi-Conversation Task Evaluate selective Auditory Attention Decoding (AAD). Two parallel stories with continuous context & turn-taking. Two independent HRTF sources moving dynamically within ±90°. Diotic pedestrian/babble noise (-9, -12 dB) 56 trials
(~45 min)
✓ EEG & Physio
✓ Behavior & Audio

Recording Modalities

All signals were synchronously recorded and streamed through g.tec HIamp and Simulink GUI at 1200 Hz, applied with 60 Hz notch filter:

Modality Specification Description
EEG 64 channels, 1200 Hz High-density neural recordings (g.tec g.HIamp)
Pupil Dilation Binocular, 60 Hz (Unified to 1200 Hz) Pupil diameter measurements (Tobii Pro Nano)
Gaze Location X and Y axes, 60 Hz (Unified to 1200 Hz) Screen-coordinate gaze tracking mapped to screen size (Tobii Pro Nano)
Respiration Flow 1200 Hz Nasal airflow monitoring
Respiration Effort 1200 Hz Thoracic expansion belt tracking
GSR 1200 Hz Galvanic skin response
PPG 1200 Hz Photoplethysmography raw optical signal
Heart Rate 1200 Hz Derived beat-by-beat heart rate from PPG
SpO₂ 1200 Hz Peripheral oxygen saturation
Temperature 1200 Hz Skin temperature
Accelerometer Triaxial, 1200 Hz Head motion tracking

Participant Information

  • Sample Size: 50 participants
  • Age Range: 18-38 years (Mean = 24, SD = 4.5)
  • Sex at Birth: 28 male, 22 female
  • Inclusion Criteria: Self-reported normal hearing, no history of neurological or attention disorders
  • Ethics: IRB approved, written informed consent obtained, participants compensated
  • Data Collection Period: October 2022 - June 2025

Data Structure & Download

Dataset Structure

The final dataset structure will be finalized and documented here upon release.

MOV-AAD/
├── participants/
│   ├── sub-001/
│   │   ├── eeg/
│   │   ├── physio/
│   │   ├── behavioral/
│   │   └── stimuli/
│   ├── sub-002/
│   └── ...sub-050/
├── stimuli/
├── preprocessing_scripts/
├── README.md
└── dataset_info.json

File Formats

  • EEG data: 64 channels @ 1200 Hz with 60 Hz notch filter
  • Physiological signals: 1200 Hz sampling rate
  • Audio recordings: Synchronized playback recordings for temporal alignment
  • Behavioral data: Event markers and timestamps
  • Spatial information: Azimuth trajectories

Download Instructions

📦 Via Zenodo

Download the complete dataset from Zenodo:

Download from Zenodo (Available Upon Publication)

DOI: [Will be added upon publication]

Dataset will be released upon paper acceptance

💻 Via Command Line

Download using wget or curl (available after publication):

# Using wget
wget [ZENODO_URL]

# Using curl
curl -O [ZENODO_URL]

# Extract
unzip MOV-AAD-dataset.zip

🐍 Via Python

Example code for loading the dataset:

# Example loading code will be provided
# upon dataset release with full documentation

Preprocessing Notes

  • EEG channels with abnormal variance are automatically detected and replaced via spherical interpolation
  • 3s pre-onset baseline segment included for each trial
  • Temporal alignment established via cross-correlation of recorded audio with original stimulus
  • Peripheral modalities provided in minimally processed form (1200 Hz with hardware notch filtering)
  • Preprocessing scripts will be provided for reproducibility

Baseline Results

Experimental Setup

We provide baseline results using standard decoding approaches to facilitate comparison with future work. All analyses use leave-one-trial-out cross-validation.

Models Evaluated

  • Envelope Reconstruction: Backward model (mTRF Toolbox v2.3) with 1-8 Hz band-pass filtered EEG, 0-400ms time lags
  • Trajectory Reconstruction: Delta band (0.05-2 Hz) and Alpha band (8-12 Hz) spatial trajectory decoding, 0-250ms lags
  • Auditory Attention Decoding (AAD): Linear Discriminant Analysis (LDA) using reconstruction correlations as features
  • Intersubject Correlation (ISC): Correlated component analysis across participants for stimulus-locked responses

Citation

If you use the MOV-AAD dataset or baseline models in your research, please cite our paper:

@article{mov_aad_2026,
  title={MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention During Moving Conversations},
  author={LastName, FirstName and Collaborator, Gavin and JointAuthor, Author},
  journal={arXiv preprint arXiv:XXXX.XXXXX},
  year={2026}
}
🔎 View on Google Scholar