Cross-linguistic vowel category formation: a comprehensive study of formant dynamics and intelligibility in Arabic and English vowel systems
Журнал: Вестник Южно-Уральского государственного университета. Серия: Лингвистика @vestnik-susu-linguistics
Рубрика: Взаимодействие языков в искусственной языковой среде
Статья в выпуске: 2 т.23, 2026 года.
Бесплатный доступ
This study explores vowel inherent spectral change and intelligibility across languages, focusing on differences between Arabic and English vowel systems. The aim of the research is to identify phonetic variation in cross-linguistic vowel acoustics that will provide the basis for Cross-Linguistic Vowel Category Formation. The paper scientific novelty is determined by obtaining new evidence that vowel category formation and intelligibility are shaped by language-specific weighting of either spectral change or duration. Static and dynamic formant properties were assessed to evaluate their contribution to the formation of vowel categories and resulting perceptual robustness across the two languages. Cross-language identification performance was significantly poorer (Arabic listeners on English vowels: 72%; English listeners on Arabic vowels: 70%), with errors characterized by static overlap and loss of dynamic contrast. Trajectory metrics were strong predictors of noise robustness, especially for English vowels. The results provide novel evidence for the expansion of frameworks accounting for cross-linguistic phonetic interaction by integrating dynamic formant cues into the Speech Learning Model (and Perceptual Assimilation Model) model framework.
Короткий адрес: https://sciup.org/147254736
IDS: 147254736 | УДК: 81’33 | DOI: 10.14529/ling260205
Формирование межъязыковой категории гласных: комплексное исследование динамики формант и разборчивости в системах арабских и английских гласных
Данное исследование сфокусировано на анализе спектральных изменений гласных, влияющих на разборчивость речи, на примере арабского и английского языков. В ходе анализа установлены фонетические различия в акустике гласных двух языков, которые могут быть основой для формирования межъязыковой категории гласных. Проанализированы статические и динамические свойства формантов и определена их роль в формировании категорий гласных и их восприятие в двух языках. В сравнительном аспекте идентификация гласных отмечена достаточно низкими показателями (арабские студенты, изучающие английский язык, корректно различают 72 % английских гласных, для английских студентов, изучающих арабский язык, этот показатель составляет 70 %). Основные ошибки заключаются в статическом наложении и потере динамического контраста. Показатели траектории были сильными предикторами устойчивости к шуму, особенно для английских гласных. Полученные результаты свидетельствуют о возможности интеграции динамических формантных сигналов в модель обучения иноязычной речи и модель перцептивной ассимиляции.
Текст научной статьи Cross-linguistic vowel category formation: a comprehensive study of formant dynamics and intelligibility in Arabic and English vowel systems
The differences between Arabic and English vowel systems in inventory size, spectral properties, and temporal characteristics are significant enough to influence how their speaker learners categorize vowels, which means that these aspects themselves or as factors for L1-L2 influence or bidirectional comparison become two of the most critical aspects. The modern standard Arabic language and its major dialects use a three- to five-vowel quality system with tiny differences in length (/i, iː, a, aː, u, uː/ (and perhaps /eː/ or /oː/)), whereas differences in formant dynamics contribute to the 10–12 monophthongal and diphthongal phonemes of General American English [15]. Understanding these differences is rooted in theoretical models of cross-linguistic speech perception. The edited Speech Learning Model (SLM-r): L2 phonetic categorization is reliant upon perception of how phonetically diverse L1 sounds are from their L2 counterparts, how accurate representations of early L1 categories are formed and the frequency and variety of L2 stimuli [14, p. 3), although space remains for mechanisms for developing both sets over one’s life span. In parallel, the Perceptual Assimilation Model (PAM) relates discrimination difficulties in nonnative speech to their assimilation into native phonological categories [9]. The combination of these frameworks forms our theoretical background to describe the formation of vowel categories influenced by static and dynamic acoustic cues. Another important acoustic dimension that goes beyond static formant targets is vowel inherent spectral change (VISC) – the change in the frequencies of formants that is systematic and intrinsic to vowel identity, even within traditionally monophthongal vowels [22, 24]. Indeed, dynamic formant properties – namely, trajectory length, velocity and direction – have been shown to provide a substantial contribution to the identification or intelligibility of vowels (with most classification tasks showing better performance utilizing these than static midpoint measures) [15, p. 3105–3108; 16]. We elicited productions from 80 participants (20 Arabic monolinguals, 20 English monolinguals, and 40 Arabic-English bilinguals) with /hVd/-like carrier contexts. Both static midpoint values and dynamic parameters (trajectory length, velocity, vector length, and functional PCA) were used to analyse formant trajectories (F1–F3). Listeners from both language backgrounds performed forced-choice vowel identification tasks in quiet and noise (+5 dB SNR) to assess intelligibility.
The research tasks are the following:
-
1. to study dimensions that characterize the inherent spectral change of a vowel focusing on differences between Arabic and English vowel systems.
-
2. to measure the perceptual-level bottom-up acoustic-perceptual linkage for English in both native (quiet) and noise.
-
3. to apply knowledge of differences between Arabic and English vowel systems for second-language pronunciation pedagogy.
The research methods selected were because of the overall goal of the study and the specific research tasks set forth. For the first task an experimental method to data collection was adopted, specifically for acoustic data associated with phonetically controlled speech materials. Descriptive and statistical analyses of formant trajectories, and resonant frequency paths were also performed to allow for in-depth examinations of dynamic spectral patterns (Thin Vowel Inherent Spectral Change, or Thin VISC) which could be examined between the two languages. We designed and ran perceptual experiments of native (quiet) and noise conditions for English to evaluate the acoustic–perceptual link at the perceptual level (second task). We systematically compared listeners' performance for dynamic versus static acoustic cues, with stimuli either presented in quiet or intermixed across a range of signal-to-noise ratios. The integration of acoustic and perceptual backgrounds thus established the robustness with which dynamic spectral information is used in vowel perception. To inform the third task production data from Arabic-English bilingual speakers was collected and analyzed in relation to theoretical predictions derived from the revised Speech Learning Model (SLM-r). Interpretation of the findings was framed in terms of L1 category formation constraints and their effects on L2 compressed VISC vowel category precision. It enabled concrete pedagogical recommendations to be derived based on empirical evidence from acoustic and production research.
The research materials are based on the phonetic frameworks adopted in vowel analysis. We focus on the dynamics of formants and their influence on the formation of the vowel category and intelligibility using the example of Arabic and English speech in order to find out to what extent subtle static and dynamic information differently affect the formation of interlanguage differences and L1–L2 interactions in various regional variants.
Practical value of the study is determined by the fact that it offers practical advice for second-language pronunciation pedagogy, particularly in regard to designing effective targeted training for Arabic-L1 English learners by specifying the differential contribution of static versus dynamic cues are suggested. It also has practical significance for developing speech recognition systems that adapt to accents, and theoretical models of cross-linguistic phonetic learning.
Literature Review
The theoretical background of the research was the works of [2, 8, 9, 11, 13, 14, 15, 20] which examines theoretical models of perception of interlanguage speech. The study of the static and dynamic properties of formants, as well as their significance for identifying vowels, is reflected in the works [3, 16, 17, 18, 22, 24, 26, 27]. These studies collectively provide theoretical scaffolding Cross-Linguistic Vowel Category Formation.
In English, classic accounts of how vowels are specified have emphasized static measures that characterize the vowel over what is presumed to be a steady-state portion of the vowel (e.g., [18, p. 245]), such as mid-point formant frequencies (F1 and F2). Indeed, these static analyses have been central to mapping the vowel space, as with studies of American English vowels where mid-point formant values categorize phonemes like /i/ versus /ɪ/ [15, p. 3100– 3102). However, recent advances show the limitations of static methods and present difficulties in maintaining both dialectal variation and contextual effects, underscoring the case for dynamic models which incorporate vowel inherent spectral change [VISC], i.e., the systematic formant trajectories across steady-state objects’ [24, p. 118). In contrast to vowels are more typically discriminated against by temporal features e.g. path length, speed and direction of the dynamic specification that could account for their similarly low dynamic visual-modality speechreading cues. For instance, using polynomial fits generated for dynamically calculated formant trajectories instead of static midpoints has proved a better predictor of vowel category discrimination for at least one dialect (West Australian English) than has the use of static midpoints, and especially with respect to the lax vowels [10, p. 3]; another, cross-dialectal study of American English dialect data also suggests that dynamically-calibrated shifts in formants is more useful than a simple static measure (i.e., both in /aɪ/ or /aʊ/) in accounting for regional variation via comparisons to digital archive material [26, p. 2]. Similarly, with respect to British English dynamic properties of a vowel-like formant (e. g., velocity) also offer perceptual robustness as static representations tend systematically to underestimate the aforementioned spectral variability [25, p. 11]. Comparative work in English varieties emphasizes the connection: in clear speech dynamic movement of formants aids vowel contrast, with greater extent of trajectory for tense [21, p. 67]. Nearey [23] explores relational properties in vowel perception, providing a review of dynamic cues, and how they may form a complementary strategy with static ones to facilitate accurate identification, noting that “[i]n noise-degraded environments an enhancement of the accuracy of our identification is achieved by relaying on formant contours” [23, p. 2085–2086]. This binary between static and dynamic is of course not immutable; hybrids exist, where midpoint values are coupled with (or supplanted by) measures taken along the trajectory to allow for a more complete picture, e.g. in cross-dialectal evaluations demonstrating that dynamic information has greater salience in eliminating overlap between vowel spaces [18, p. 247).
Arabic-English vowel systems In this line of inquiry, Arabic-English vowel systems across languages have been another area of focus as a perceptual and productive challenge arising from conflicting inventories [14, p. 20], in respect to the smaller length-based vs larger feature based and dynamic system that is English. According to Evans and Alshangiti [7, p. 7] Arab learner ordered that as the English L2 vowel occupy an Arabic category which rely on reduced much lot may also perhaps this account some kind of variation identification accuracy consonants brim top on correctness right than/æ-ʌ/or /ɛ-ɪ/on Saudi load reach. In contrast, English-Arabic bilinguals acquire their L2 Arabic gradually while the underlying substrate does not change (i.e., L1 English is preserved). Cross-language differences in vowel intelligibility find acoustic correlates under degraded (i.e., noisy) conditions. In environments where multiple languages coexist, the intelligibility of English vowels in noise is higher than that of Arabic or Mandarin syllables due to their larger vowel spaces. Intelligibility suffers from Arabic-English interference; specific adaptations of speakers' VOT show category shifts for Saudi English, and listeners perceive vowel assimilation in noise [13, p. 127). Therefore, more extended correlates of vowels are associated with clearer speech: larger formant ranges lead to improved vowel intelligibility in hearing-impaired listeners [12, p. 2].
Further focus is on stating how the procedures were for examining both formant dynamics, and formation of vowel categories and intelligibility that exist across either Arabic or English vowel systems. This design combines a detailed examination of acoustic production with perceptual assessment of intelligibility, using protocols used for other cross-linguistic studies of vowels conducted by Arabic L1 speakers for English vowels.
The research collected data from three participant groups, which allowed for both monolingual baselines and cross-linguistic comparisons:
– Native Arabic speakers (L1 Arabic monolinguals): Twenty participants (10 female, 10 male), aged 20–35 years (range M = 27.4, SD = 4.2), were native speakers of Hijazi or Palestinian Arabic dialect, with no reported history of speech/hearing disorders or long-term English exposure. Four of these speakers produced Arabic vowels, and eight served as listeners in intelligibility tasks for Arabic vowels. L1 English monolinguals:
– 20 native English speakers (10 female, 10 male), aged 22–38 years (M = 28.1, SD = 5.0), combined from General American and Standard Southern British English varieties, recruited to establish a baseline for both productions of the target set of English vowels as well as intelligibility measures for those sounds.
– L1 Arabic → L2 English bilinguals: 40, aged 18–35 years (20 female, 20 male; M = 26.8, SD = 4.5); late learners with intermediate to advanced proficiency (self-reported CEFR B2-C1 levels; for IELTS/TOEFL equivalents where available; mean length of exposure to English Language lessons: 8–12 years). This group produced Arabic and English vowels to permit examination of L1–L2 interactions.
All participants were right-handed, had normal or corrected-to-normal vision/hearing and written informed consent was obtained. Sample sizes are consistent with other studies employing acoustic measurements of Arabic-background language, specifically comparing individuals across a range of vowel production tasks (n=20–65 speakers per group) and perceptual tasks that pertain to both languages using listener pools within the same ranges.
Stimuli were maximally controlled in phonetic context to limit coarticulatory effects, allowing normal production:
– Target vowels: For English, the 11–12 total of monophthongal vowels present in General American or SSBE (e.g., /i, ɪ, e, ɛ, æ, ɑ/ɔ/oʊ/ʌ/ɜ/); for Arabic ones core long-short pairs quota (/i iː/, where dialect-dialect manner inclusions – potentially including native assimilable cycles like /(ma)/-, along with common bunches).
– Carrier phrase: Vowels elicited in /hVd/ (English), or identical CVC/hVd/-style frames (Arabic, phonotactically parsed per e.g., adjusting so that h is not adjacent to the vowel by replacing it with appropriate bilabials like wd i.e. using /bVd/ or /hVd/, but there is no example formation possible in
Arabic+so used for English only) This was given two tokens per speaker for each of the five vowels, resulting in roughly 200–400 tokens per language/group.
– To record, participants read randomized lists in a quiet room using a head-mounted microphone (e.g., Shure SM10A) attached to a digital recorder (44.1 kHz; 16 bits). Order effects were controlled by counterbalancing Arabic and English sessions.
First, formant Extraction; Praat function [18, p. 245], was applied in order to extract F1, F2 and F3 after tracking the formants. Burg algorithm parameters: time step 0.00625 s; maximum formants 5; ceiling 5000 Hz (males)/5500 Hz (females); window length 0.025 s and pre-emphasis of 50 Hz; Tracks were visually checked for tracking errors (e.g., octave jumps) and corrected when needed.
Secondly, static Measures: Formant values at the vowel half-life (50% duration) were used as static target ranges. Second, metrics including vowel space area (convex hull of F1–F2 midpoints) and dispersion (Euclidean distances from centroid). The used speaker normalization, which either used Lobanov z-scores or Bark-transformed difference to remove the interspeaker variation of on various speech features.
Thirdly, Dynamic Properties: Vowel-inherent spectral change (VISC) was measured using: - Trajectory length (TL): Euclidean distance travelled along path in F1–F2 space at normalized time points (e.g. 11–13 points at 0–100% duration in intervals of 10%). - Formant velocity (FV): Change rate of H1/F2 (ΔF1/Δt, ΔF2/Δt) Vector length (VL): Onset-to-offset displacement Functional Principal Component Analysis: for timeseries formant trajectories. Other features: vowel duration, F0 median and F3 mean value on each euclidean path moved during angular displacement. These analyses applied multi-method logic: group distinctions using linear mixed-effects models (lmer in R), followed by discriminative analysis (quadratic, leave-one-out cross-validation) of static vs. dynamic model classification accuracy.
Identifying vowels: Isolated real vowel tokens (from production recordings) presented in quiet & noise (+5 dB SNR babble; via headphones): perceptual tasks include identifying cross-language intelligibility They focused largely on the phonology of particular languages, and were driven by specific challenges posed to them for solving behavior (e.g. Poor-vowels-signal, and other perceptual studies) Auditory perception genre n = 20/group in across-language-listeners (80 ratings /1; that is Arabic listeners rating English vowels; English listeners judging Arabic vowels; bilinguals judging both) – Forced-choice identification a primary measure used such as present here to observe assimilation – Phoneme-based – Confusion matrices for cross-linguistic assimilation mapping More measures; Sentence-in-noise intelligibility (set/potpourri of subset of vowels in carrier sentences), adaptive SNR to assess 50% thresholds.
Data were analysed in R (R Core Team, 2023): Linear mixed-effects models (lme4 package) for all formant/duration differences, with fixed effects (language/group/vowel/sex) and random intercepts/slopes (speaker/vowel token). Vowel space comparisons: Multivariate analyses (MANOVA) Discriminant analysis and classification models (MASS package) to evaluate contributions of the cues (static, dynamic, combined). Correlation/regression for acoustic-perceptual associations (e.g. trajectory metrics predicting identification accuracy) Statistical significance level α = 0.05 (Bonferroni corrections for multiple comparisons applied where appropriate) Effect size statistics (partial η²; Cohen's d) are reported for significant contrasts. This approach provides a solid, replicable examination of formant dynamics and intelligibility across Arabic and English vowels targets. 46 web pages Explain in detail fPCA Compare French vowel systems Make it more concise.
Analyses were conducted as described in the methods, employing linear mixed-effects models (LMEMs) for group and vowel comparisons where fixed effects comprised of language (Arabic vs. English), group (monolingual Arabic, monolingual English or bilingual Arabic-English), sex (male vs. female) and vowel identity and random effects consisted of speaker and token intercepts/slopes. Effect sizes are reported for as partial η² for MANOVAs and Cohen's d for pairwise contrasts. Discriminant analyses used quadratic classifiers with leave-one-out cross-validation. Tracking artifacts accounted for 2.1% of the data, resulting in a final sample of 12,480 vowel tokens (≈156 per speaker) from 80 participants.
Static midpoint formants (F1 and F2, Lobanov-normalized) showed that the 3D arrangements of Arabic and English vowel systems differ from each other (Table 1), where more crowded categories were observed in an English system where a greater dispersion had been observed. The total vowel space area (convex hull in F1-F2 plane) was greater for monolingual English speakers (M = 4.82 Bark², SD = 0.91) than for monolingual Arabic speakers (M = 3.15 Bark², SD = 0.72; LMEM: β = 1.67, SE=0.28, t =5.96, p <.001). 001, d = 1.12). Intermediate values were found for bilinguals (M = 3.98 Bark², SD = 0.85), indicative of partial L2 expansion and conservative L1 compression (group contrast: F(2,77) = 14.23, p < .001, η² = .27).
Midpoint F1 and F2 values (in Hz, unnormalized for authenticity) are summarized in Table 1 (averaged across both sexes; males were generally 10–15% lower than females due to the length of their vocal tracts). English vowels fell further out into space, with high-front /i/ at F1 = 312 Hz, F2 = 2298 Hz – versus Arabic /iː/ (F1 = 348 Hz, F2 = 2156 Hz). There was mid-central overlap across languages: English /ʌ/ (F1 = 612 Hz, F2 = 1387 Hz) overlapped Arabic short /a/ (F1 = 598 Hz, F2 = 1421 Hz), and dispersion metrics further indicated that the whole vowel-space of English was more variable than that of Arabic (SD F1 = 45 Hz vs. Arabic SD F1 = 32Hz; [MATH]). 003, η² =. 11).
This is evaluated against monolingual /ɛ/ and /æ/, which were acoustically farther away from each other (d = 1.45; d = 0.82) in bilingual productions of English than they were in monolingual variety productions (t(38) = 3.12, p =. 004). Sex effects were consistent across measures, with females exhibiting higher frequencies of F2 (β = 124 Hz, SE = 18, t = 6.89, p <. 001).
VISC was greater in the English vowels from dynamic analyses, with reliance on durational cues more pronounced in Arabic. Trajectory metrics were calculated at 11 time normalized points (0–100% duration).
At the group-level, English's vowels produced longer trajectory lengths than did Arabic (LMEM: β = 1.13, SE = 0.19, t(158) = 5.95; p < .001): M[English]TL = 2.41 Bark (SD = 0.57), M[Arabic]tl = 1.28Bark (SD = 0.41). 001, d = 1.08). Mean FV in English was greater than for Arabic (M FV: ΔF2/Δt = 4.2 Hz/ms vs. M – Arabic = 2.1 Hz/ms; F(1,78) = 18.64, p <. 001, η² = 0.19). While bilinguals showed near English distributions in their L2 productions (TL M = 2.12 Bark), they maintained Arabic-like stability of the L1 compared to it (TL M = 1.42 Bark) as would be anticipated from asymmetric influence. The Arabic spectral extent of change (vector length, VL) was correlated with duration (r = .68, p <. 001) (though with F2 direction in English r = .52, p = .002). fPCA showed two principal components: PC1 (62% variance) relating to overall trajectory curvature which was greater for English; and PC2 (21%) relating to onset-offset divergence, which had increased values for bilingual English.
HCM analyses for specific vowels: HCs 3/4: English lax v/c (e.g., /ɪ, ʊ/) showed at least
Table 1
Mean Midpoint Formant Values (F1/F2 in Hz) for Selected Vowels by Group
|
Vowel |
Monolingual Arabic |
Monolingual English |
Bilingual (Arabic) |
Bilingual (English) |
|
/i(ː)/ |
348 / 2156 |
312 / 2298 |
335 / 2204 |
320 / 2267 |
|
/a(ː)/ |
598 / 1421 (short); 712 / 1298 (long) |
685 / 1325 (/ɑ/) |
612 / 1389 (short); 698 / 1312 (long) |
672 / 1341 (/ɑ/) |
|
/u(ː)/ |
352 / 1024 |
318 / 968 |
340 / 998 |
325 / 982 |
|
/ɛ/ (Eng) or equiv. |
N/A |
512 / 1845 |
N/A |
498 / 1812 |
|
/æ/ (Eng) |
N/A |
678 / 1623 |
N/A |
652 / 1589 |
unsurprising proximity to their own tensioned proximal counterparts given all co-opted 'VISC-constructs' (this is all VFC-significant and VLRA-significant): so that indicated surface number-line homogeneity evident (F two sonances clearly rising from Talus onset 1987 Hz while end off projections reached just above final Bark limit along through HNC) reaching examples. In contrast, the Arabic short /i/ was virtually a static vowel (onset vs. offset F2 = 2111 Hz vs. 2175 Hz; TL = 0.9 Bark; vowel × language: F(10,780) = 7.82, p < .001, η² = .09). Bilinguals realized English /ɪ/ with decreased divergence (TL = 2.1 Bark, more Arabic-like), resulting in merger-type patterns (confusion with /i/ in perceptual tasks). For low vowels, there was significant backing of English /æ/ (F2 value decrease: F2 onset 1689 Hz at offset 1542 Hz; angle direction -12°) and no such change for Arabic /a/ (onset range 1408 Hz to offset 1434 Hz; d = 1.32 to TL contrast). Thus, spring /aː/ could be produced long (M = 185 ms) yet not possess VISC (VL = 0.7 Bark) – in contrast to English /ɑ/ (M duration=142 ms, VL=1.6 Bark). A most unusual cross-linguistic contrast emerged for high-back vowels: the trajectory of English /ʊ/ involved centralization (F2 increase 8%), which was absent in Arabic /u/ (difference in F2 < 2%; t(78) = 4.56, p < .001). Formant extracts (from an exemplar male monolingual English speaker, token /hɪd/: Time-normalised F2 trajectory: [0%: 2012 Hz, 20%: 2124 Hz, 40%: 2208 Hz, 60%: 2289 Hz, 80%: 2321 Hz, 100% 2356Hz]; TL = 295 Bark. Arabic/hid/(mono male): [0%: 2134 Hz, 20%: 2148 Hz, 40%: 2156 Hz, 60%: 2162 Hz, 80%. 2165 Hz, 100%. 2169Hz]; TL=0.48Bark.
Also, there were detected in bilinguals, but the non L1 Arabic vowel spaces had also become compressed toward their corresponding centroids (centroid shift: 0.85 Bark from monolingual English; MANOVA: F (2,38) = 11.45, p <. 001, η² =. 38), with lower front vowel dispersion (e.g. /i-ɪ-ɛ-æ/ chain ↓18%). Reviewing studies of this type, a slight increase (0.32 Bark) in L1 Arabic for long vowels was found for bilinguals, indicating limited cultural loss. Trajectory patterns were recategorised: The mixed production in English weakly activated hybrid dynamic, and the VISC was 25% low with no other language than English (β = –0.62, SE = 0.21, t = –2.95, p =) 004), which was in close accordance (r =.45, p = .01; high TL with high IELT score).
Discriminant analyses compared classification accuracy using static (midpoint F1/F2), dynamic (TL, FV, VL, direction) and combined models. For monolingual English, dynamic cues alone yielded 82.4% accuracy (vs. static 71.6%; χ²(1) = 14.2, p < .001), rising to 91.7% combined. In monolingual Arabic, static cues were adequate (78.9% accuracy), where dynamics improved performance minimally (combined 81.3%; χ²(1) = 2.1, p =. 15). English bilingualism privileged dynamics (79.2% vs. static 68.4%; d = 0.92), while Arabic bilingualists showed a preference for the static (76.5%). Stepwise regression for cue weighting suggested that, altogether, dynamic features accounted for 48% of English discrimination variance (R² =. 48, p <. 001), compared to 22% in Arabic (R² =. 22, p =. 002). Cross-validation established robustness, with dynamic models curtailing errors (e.g., English /ɛ-æ/ misclassifications decreased 35%, across overlapping categories).
Vowel-identification accuracy was greater for monolingual listeners judging native than crosslanguage vowels (English: 88.2%, SD = 6.1; Arabic: 84.7%, SD = 5.8 vs Arabic listeners on English: 72.4%; English on Arabic: 69.8%; group × language interaction of F(2,57) = 16.8, p < .001, η² =. 36). In noise (+5 dB SNR), performance dropped more for English (71.5%) than Arabic (76.2%), which is due to the dependence on VISC (correlation TL and accuracy drop r = –.58, p < .001). Patterns of confusion: Arabic listeners assimilated English /ɪ/ to /i/ (32% errors) and /æ/ to /a/ (28%), indicative of static overlap. Listeners of English confused short Arabic /a/ with /ʌ/ (25%). Bilinguals made significantly more accurate cross-intelligible judgments (L2 English: 81.6%), with dynamics predicting accuracy (regression: β_TL = 0.41, SE = 0.12, t = 3.42, p = .001). Trajectory metrics that were negatively correlated with sentence-in-noise thresholds (r = -0. 49 for VL, p =. 003), emphasizing the strengthening effect of dynamic cues on robustness. These results further support the notion of cross-linguistic variation in the degree to which formant/c2 is relied on, as dynamics contributes to categorization and intelligibility among English listeners.
Cross-lingually, we found L1 Arabic speakers in the bilingual group produced English vowel with lesser degrees of VISC (e.g., shorter TL and lower FV for lax vowels [e.g., /ɪ/ and /ʊ/] post-education) which resulted in compression of your vowel-space and tendencies toward merging in certain mid-regions (/ɛ– æ/, /ɪ–i/). This pattern is consistent with the recent revised Speech Learning Model (SLM-r, Flege & Bohn, 2021), which proposes that new L2 phonetic categories develop based on perceived dissimilarity from L1 prototypes and recognize the quality of L2 input received and precision of respective L1 categories; equivalently-classified as similar sounds lead to merged or reorganized categories compared to a not full separation [14, p. 20–25, 45–50). In this case, the assimilation of English dynamic vowels to more static (Arabic) patterns is an example of systemlevel reorganization, bringing SLM-r predictions into spectral-temporal geometry that extends beyond static targets – it is not limited only to Yi's finding that adolescent Palestinian Arabic learners of English showed centralized, flattened trajectories with diminished modulation [5]. There were limited bidirectional effects: L1 Arabic vowels of bilinguals showed a slight expansion, suggesting some resistance to attrition at the level of experienced learners and thereby in line with patterns of high-proficiency relative maintenance common in older lifelong bilingual speakers [20].
These results thus extend models of speech learning within a theoretical space as well, with dynamic specification emerging as a dimension that captures L1–L2 interaction: target approximation may be implicated in per SLM-r equivalence classification (as proposed by SLM-r), not least along other dimensions such as re-organization of trajectory patterns during the perceptual task, while temporaldomain correlations structure hostile assimilation mappings for PAM in ways incomplete when conceived only within static-focused accounts. As for the methodology, the multi-cue findings (static + dynamic + fPCA) were sensitive to cross-linguistic variation, but dialectal specificity (predominantly Hijazi/Palestinian Arabic) and controlled CVC contexts limit generalizability outside spontaneous speech or other dialects (i.e., differing emphasis effects might emerge in Najdi → Gulf varieties). Limitations include exclusive focus on production baselines with perceptual tasks using isolated vowels, likely to underestimate contextual influences of coarticulation; although listener groups were balanced around L1 and L2 language background, they were not representative of invariances for bidirectional bilingualism influences; proficiency controls depended at least in part on self-report/equivalent rather than consistent testing. Future work may take a longitudinal approach to track the onset of VISC, further explore dialect-sensitive moderators (e.g., emphatic versus non-emphatic contexts), include articulatory data (e.g., ultrasound) to understand trajectory origins, and evaluate dynamic cue weighting across more ecologically valid noise conditions at the sentence level.
Conclusion
The tasks indicated in the present program of research have been achieved and research findings reported here demonstrate how findings are such a fitting match with this study's tasks. The research findings showed that the phonetic dimensions associated with the inherent spectral variation in vowels (VISC) are a decisive factor in distinguishing subtle differences between speech sounds, particularly when comparing the phonetic systems of Arabic and English. It was found that these dimensions do not operate independently, but are influenced by the consonantal context, particularly when vowels occur before fricative plosives, where temporal dynamics lead to clear variations in the trajectories of resonant frequencies. Furthermore, production results from bilingual speakers revealed only a partial reorganisation of the vowel space in the second language, reflecting the influence of first-language constraints, which is consistent with the predictions of the SLM-r model.
Perceptual studies confirmed acoustic-perceptual linkage for English in both native (quiet) and noise by revealing that dynamics associated with VISC were a powerful cue for identification, outperforming static cues and conferring advantages over considerably detracting levels of noise. We further demonstrated that the cross-language perceptual asymmetries observed closely reflected MDS-derived assimilations predicted by the Perceptual Assimilation Model (PAM), supporting characterizations of native language experience with Arabic vowels as facilitating and shaping L2 English mapping and discrimination of novel L2 vowel categories.
The research translated knowledge of Arabic and English vowel system differences into teaching pronunciation for second language acquisition. Conversely, production data from bilinguals who consistently spoke both perceptions within the constraints of L1-based category formation predicted by Evans and Baker. And some Arabic-English speakers showed partial reorganizations of their L2 vowel spaces that were consistent with the predictions of this hypothesis. L1 drove precision on L2 categories – these VISC compressions make the relationship clear. This knowledge informs pedagogical practices and can help identify what must be taught (temporal-spectral dynamics needs to be directly taught in order for L2 vowels to map more accurately, accent effects need to be mitigated, and pronunciation training made possible using relative comparators based on user’s experience with nativelike English such as around Arabic speakers). To sum up, the phonetic information revealed by the experimental acoustic data collection implemented in traditional media coupled with perceptual testing and descriptive-statistical analyses of resonant frequency paths supported that dynamic spectral dimensions are essential features of a phonetically appropriate articulation of vowels (in terms of both their distinguishing senses and learning modalities). These research findings highlight the importance of temporal-spectral dynamics in models of speech perception and production which may impact broader matters like second-language pedagogy, accent-adapted speech technology, or even phonetic learning in more finetuning structure.
Future Research Directions : explore the relative roles of vowel inherent spectral change (VISC) and consonantal context, focusing on how dynamic cues differ across vowels preceding voiceless stops in Arabic compared to English, and whether such contextual modulation is subject to languagespecificity leading to cross-linguistic category formation; explore whether and to what extent VISC compression in productions of L2 English by bilingual Arabic–English speakers is also observed in perceptual learning, using longitudinal designs that capture developmental change; examine the time course of when dynamic cue use in online vowel recognition (via eye movement or EEG, for example) under different noise conditions.