Introduction: The use of brain–computer interfaces (BCIs) to decode imagined speech has significant clinical and assistive potential.
Methods: Twenty-four studies investigated covert speech decoding between 2009 and 2025 using electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), or hybrid EEG–fNIRS systems.
Results: Early research (2009–2012) primarily focused on analyzing phonemes and syllables with EEG, achieving accuracy rates around 75%. From 2013 to 2017, convolutional neural network (CNN)-based phoneme decoding produced highly variable results (40–83%), with more complex multiclass tasks occasionally performing poorly (as low as 26.7%). Since 2018, binary paradigms such as yes/no responses have reached 64–100% accuracy. CNN variants (about 83.4%), AlexNet (90.3%), and LSTM-RNNs (92.5%) demonstrated notable improvements, whereas architectures, like EEGNet and SPDNet often underperformed (24.79–66.93%). In hybrid EEG–fNIRS studies, reported accuracy ranged from approximately 53% with CNN-based decoding to 70.45±19.19% with an RLDA-based decision-level fusion approach, although direct comparisons are limited by differences in tasks, fusion strategies, and validation settings.
Conclusion: Although deep learning and multimodal systems have potential for enhancing imagined speech decoding, there are still major challenges related to generalization, variability, and robustness.
نوع مطالعه:
Review |
موضوع مقاله:
Cognitive Neuroscience دریافت: 1404/5/16 | پذیرش: 1404/10/20 | انتشار: 1404/12/10