科学・技術
EEGは脳が同時に2つの音声ストリームをエンコードできることを示す
EEG shows brain can simultaneous encode two speech streams (journals.plos.org)
要約
複数の話し手がいる騒がしい環境で注意を切り替える際の脳のメカニズムをEEGを用いて調査した研究です。注意を切り替える際、脳は古い音声ストリームから完全に離れる前に、一時的に両方の音声ストリームを同時にエンコードすることが明らかになりました。このプロセスは、α波パワーの低下と関連しており、語彙的文脈の更新も関与している可能性が示唆されています。
全文翻訳
多人数が話す状況で円滑にコミュニケーションをとるには、持続的な注意と迅速な注意の切り替えを巧みに組み合わせる必要があります。神経生理学の文献は持続的な注意の神経基盤について詳細な洞察を提供していますが、注意の切り替えがどのように行われるかについては、かなりの不確実性が残っています。本研究では、没入型の多人数環境における健聴成人からのEEG記録を使用し、背景の雑音の中で競合する2つの音声ストリームの神経エンコーディングを測定しました。参加者は15〜30秒ごとにストリーム間の注意を切り替えるように指示されました。神経追跡はTemporal Response Functions(TRF)を介して評価され、注意の焦点を確実にデコードできることが確認されました。我々の結果は、注意切り替え中の非対称な離脱および関与プロセスを示しており、新しいターゲットストリームの神経追跡が、前のターゲットから離れる前に現れることを示しています。これにより、2つの音声ストリームの一時的な同時エンコーディングが明らかになりました。この遷移はEEGのα波パワーの低下と密接に一致しており、注意切り替えの異なる段階での認知的努力について情報を提供します。次に、注意切り替え後の語彙的文脈がどのように更新されるかを判断するために、大規模言語モデル(LLM)を使用して構築された4つの文脈蓄積戦略を比較し、語彙予測メカニズムを反映する皮質活動を分離しました。我々の発見は、聴覚注意シフトの背後にある時間的および文脈的メカニズムの両方を解明し、注意を切り替えた後にリスナーが語彙的文脈のリセットを実行する可能性を示唆しています。動的な注意の再配分に焦点を当てることで、この研究は、複雑な聴覚環境における柔軟な音声処理に対する脳の能力についての洞察を提供します。
引用: Carta S, Aličković E, Zaar J, López Valdés A, Di Liberto GM (2026) Competing speech streams are simultaneously represented in the human cortex during attention switching. PLoS Biol 24(7): e3003876. https://doi.org/10.1371/journal.pbio.3003876
Academic Editor: Manuel S. Malmierca, Universidad de Salamanca, SPAIN
Received: July 3, 2025; Accepted: June 12, 2026; Published: July 16, 2026
Copyright: © 2026 Carta et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All data supporting the findings reported in this manuscript are freely accessible without restriction. The EEG pre-processed dataset, the resulting analysis files, and the analysis code are publicly available on the open repository Zenodo (https://zenodo.org/records/20569817). The EEG recordings are provided following the Continuous-event Neural Data (CND) format standard. The associated speech stimuli can also be found in the same repository, within the STIMULI folder.
Funding: S.C., A.L.V., and G.D.L. were supported by the William Demant Fonden (https://www.williamdemantfonden.dk/), under grants 21-0628 and 22-0552, and by Taighde Éireann – Research Ireland (https://www.researchireland.ie/) under grant No. 18/CRT/6223. G.D.L. additionally conducted this research with the financial support of Research Ireland at ADAPT, the Research Ireland Centre for AI-Driven Digital Content Technology (https://www.adaptcentre.ie/) at Trinity College Dublin [grant 13/RC/2106_P2]. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Abbreviations: EEG, electroencephalography; EOG, electro-oculography; EMG, electro-myography; ERSP, event-related spectral perturbation; iEEG, intra-cranial electroencephalography; FDR, false discovery rate; fMRI, functional magnetic resonance imaging; ICA, Independent Component Analysis; IQR, interquartile range; LLM, large language model; MEG, magnetoencephalography; PSD, power spectral density; RMS, root-mean-squared; SE, standard error; SEM, standard error of the mean; SNR, signal-to-noise ratio; SPL, sound pressure level; TRF, Temporal Response Functions
Introduction
To understand speech in multi-talker environments, listeners single out the target speaker from competing sound streams [1–3]. The neurophysiology of this selective attention process has been widely studied with simulated cocktail-party scenarios [4,5], shedding light on how our brains segregate a target stream from competing speech streams, and enabling the transformation of the target speech into linguistic meaning. While the extent to which masker speech streams are processed remains highly debated [6–8], there is no doubt that there are considerable differences between the processing of target and masker speech, which have been measured with various technologies, such as non-invasive electroencephalography (EEG) [1,9], intra-cranial electroencephalography (iEEG) [10], magnetoencephalography (MEG) [3,11] and functional magnetic resonance imaging (fMRI) [12,13]. That work could pinpoint precise loci in the auditory cortical areas where that segregation emerges [14] as well as measuring the substantial (but not total) suppression of linguistic processing for the masker speech [1,15–17]. However, neurophysiology literature in this field has almost entirely focused on sustained attention tasks [2,10], leaving considerable uncertainty on the neural underpinnings of attention switching. Dynamic switching paradigms have been widely used in the domain of cognitive control studies to probe for cognitive flexibility and cognitive stability [18]. In those experiments, participants are often required to flexibly adapt their behavioral response depending on new instructions, initiating a task-switch [19–21]. For example, given a single digit, they are required to classify it either based on parity, i.e., whether it is even or odd, or based on relative magnitude, i.e., whether the digit is greater than or less than 5 [22]. In these paradigms, the switch-cost is the increase in reaction time or error rate when switching from one task to the other. Similar behavioral paradigms have also involved simple speech stimuli in multi-talker settings [23–25]. However, the main interest of those tightly controlled experiments was to model the process of target speech selection as one particular instance of a task-switching problem, i.e., target stream selection could either depend on spatial location or voice identity [23], rather than focusing on the dynamic aspect of attention re-allocation per se in naturalistic multi-talker scenarios. As such, very little is known on how a flexible reorienting of attention might impact speech processing of continuous competing streams. In recent speech neurophysiology research, experimental paradigms have started to include switches of attention as a tool towards tailored EEG/MEG methodological advances in the domain of attention decoding [26,27], or to investigate how sustained speech attention unfolds for moving auditory objects [28]. However, to the best of our knowledge, only one previous study has specifically focused on the neurophysiology of attention switching in multi-talker scenarios, relating the neural encoding of speech during attentional re-orienting with EEG alpha activity and pupil dilation dynamics [29]. Those findings proved that the neurophysiology of attention switching can be studied non-invasively. Building on that work, our study sheds light on the exact neural dynamics supporting the steering of attention between two competing speech streams, disengaging from the previous target stream while engaging to the new one. In this study, we measure the neural encoding of speech using a range of encoding window lengths, as listeners steer their attention from one speaker to another.