Phonetics for English Teachers

This section summarizes phonetics information necessary for English teachers when conducting pronunciation instruction. Focusing on comparison between English and Japanese sound systems, it explains teaching points that are particularly helpful for teachers whose native language is Japanese.

A teacher giving pronunciation instruction in a classroom
1

Pronunciation model and goal

A pronunciation model refers to the pronunciation that students hear and use as a reference in the classroom. To ensure that students are not confused, it is best to select a single accent as the model for imitation. In Japanese English education, General American (the standard American pronunciation) is commonly employed. In contrast, a pronunciation goal refers to the level of pronunciation that individual learners aim to achieve. While some learners aspire to sound exactly like a native speaker, others believe it is sufficient to acquire pronunciation that is intelligible enough to get by while traveling; thus, goals vary from learner to learner. Therefore, it is important to distinguish between the pronunciation model and goal.

Pronunciation model

||

the pronunciation that students hear
and use as a reference

Pronunciation goal

||

the pronunciation level that individual
learners aim to achieve

2

Diversity of English speakers and their pronunciation

The use of English, widely spoken as both an international language and a lingua franca, can be categorized into three groups [1] Kachru, B. B. (1992). Teaching world Englishes. In B. B. Kachru (Ed.),The other tongue: English across cultures (2nd ed.) (pp. 355-365). Springer. . The first group comprises countries and regions known as the Inner Circle, where English is spoken as a native language, for example, the United States, Canada, the United Kingdom, Ireland, Australia, and New Zealand. The second group includes countries and regions known as the Outer Circle, where English is used daily and often designated as one of the official languages, for example, Singapore, India, the Philippines, Jamaica, and Nigeria. The third group, the Expanding Circle, includes countries and regions where English is studied as a foreign language in schools, for example, Japan, China, South Korea, France, Germany, Saudi Arabia, and Brazil. A defining characteristic of English is that non-native speakers significantly outnumber native speakers.

  • Inner Circle

    the United States, Canada,
    the United Kingdom, Ireland,
    Australia, New Zealand

  • Outer Circle

    Singapore, India, the Philippines,
    Jamaica, Nigeria, Kenya

  • Expanding Circle

    Japan, China, South Korea,
    France, Germany, Saudi Arabia,
    Brazil, Egypt

3

Intelligible and comprehensible pronunciation

People having a conversation in English
People having a conversation in English

Non-native speakers of English often retain traces of their native language, which is referred to as “accentedness.” While eliminating these traces is necessary if one aims for native-level pronunciation, there exist cases where pronunciation remains sufficiently intelligible even with such traces [2] Derwing, T. M., & Munro, M. J. (1997). Accent, intelligibility, and comprehensibility: Evidence from four L1s.Studies in Second Language Acquisition, 19, 1–16. . Currently, the goal for learners is not pronunciation devoid of traces of their native language, but rather pronunciation with significant intelligibility and comprehensibility. “Intelligibility” refers to the extent to which the listener understands what the speaker intended to convey. “Comprehensibility” refers to the amount of effort required by the listener to understand [3] Derwing, T. M. and M. J. Munro. (2015). Pronunciation fundamentals. John Benjamins Publishing Company. . The pronunciation learners should aim for is not “pronunciation equivalent to that of a native speaker,” but rather “pronunciation that is intelligible and comprehensible” [4] Levis, J. M. (2005). Changing contexts and shifting paradigms in pronunciation teaching. TESOL Quarterly, 39(3), 369-377. [5] Levis, J. M. (2020). Revisiting the intelligibility and nativeness principles. Journal of Second Language Pronunciation, 6(3), 310-328. .

4

Knowledge and skills necessary for pronunciation instruction

To effectively provide pronunciation instruction, teachers require knowledge and skills in three key areas [6] 杉本淳子・内田洋子 (2020).「英語教員養成における音声学教育:日本人英語教員のための〈教職音声学〉試案」『音声研究』24, 22-35. . The first is “practical skills of pronunciation and listening.” Teachers must develop intelligible and comprehensible pronunciation that serves as an effective model for students, as well as the ability to understand a variety of English accents. The second is basic “knowledge of phonetics.” By understanding the differences between the learner’s native language and the English sound system, teachers can ascertain the pronunciation challenges experienced by learners. The third, “knowledge and skills of pronunciation instruction,” includes the ability to explain the pronunciation mechanism succinctly to students, evaluate learners’ pronunciation, provide appropriate feedback and advice, and design pronunciation activities tailored to the learners’ level.

Three key areas of pronunciation instruction
5

Non-native teachers’ pronunciation goal

There are several advantages of having non-native English teachers provide pronunciation instruction to students. Particularly, a teacher with profound knowledge of the learners’ native language can appreciate the challenges they experience and provide appropriate guidance. Since the teacher’s pronunciation serves as a model for students, it is essential for teachers to master intelligible and comprehensible pronunciation. Even if a teacher’s pronunciation retains traces of their native language, it is evaluated as appropriate for an English teacher [7] Sugimoto, J., & Uchida, Y. (2018). Accentedness and acceptability ratings of Japanese English teachers’ pronunciation. Proceedings of PSLLT-9, 30-40. . Furthermore, there is a correlation between a teacher’s pronunciation ability and their positive attitudes toward teaching pronunciation; research has shown that English teachers who are confident about their own pronunciation are also more willing to teach pronunciation [8] Uchida, Y., & Sugimoto, J. (2020). Non-native English teachers’ confidence in their own pronunciation and attitudes towards teaching: A questionnaire survey in Japan. International Journal of Applied Linguistics, 30(1), 19-34. .

6

Types of pronunciation activities

There are three types of pronunciation activities: form-focused, content-focused, and balanced [9] Muller Levis, G., & Levis, J. (2016). Integrating pronunciation into listening/speaking classes. In T. Jones (Ed.) Pronunciation in the classroom: The overlooked essential(pp. 27-42). TESOL. . For example, when practicing /l/ and /r/, a form-focused activity involves repeatedly pronouncing a list of minimal pairs (e.g., light–right, glass–grass, pilot–pirate). In contrast, an activity where students record interviews with each other and then listen to the recordings to ascertain whether they are pronouncing the /l/ and /r/ sounds accurately in spontaneous speech is considered content-focused. If controls are introduced, such as specifying the words or phrases to be used in the interview, the activity becomes balanced [10] Uchida, Y., & Sugimoto, J. (2026). Developing pedagogical knowledge for English pronunciation teaching through material evaluation. Speak Out!, 74, 33-45. . Each type of activity has its own advantages and disadvantages.

  • Form-focused
    activities

  • Balanced
    activities

  • Content-focused
    activities

7

Pronunciation assessment

Someone filling in a pronunciation assessment sheet

There are various methods for evaluating pronunciation. For example, evaluations that assess pronunciation as a whole, such as “intelligibility” and “comprehensibility,” are called holistic evaluations. In contrast, evaluations that focus on specific elements, such as consonants, vowels, word stress, or tone unit boundaries, are called analytic evaluations. It is important to start by incorporating accessible methods such as a 5-point scale for “comprehensibility” when evaluating a speech (holistic evaluation) or assessing the stress patterns of new vocabulary items introduced in a specific unit (analytic evaluation) [11] Isbell, D. R., & Sakai, M. (2022). Pronunciation assessment in classroom contexts. In J. Levis, T. Derwing, & S. Sonsaat-Hegelheimer (Eds.), Second language pronunciation: Bridging the gap between research and teaching (pp. 194-214). John Wiley & Sons. [12] 常本亜希 (2025). 「評価において何をどのように捉えるべきか」シンポジウム「教員のための音声指導と評価」外国語教育メディア学会 (LET) 第64回 (2025) 年次研究大会. .

8

Functional load and prioritizing vowels and consonants to teach

Functional load refers to the power that vowels and consonants possess to distinguish meaning in a given language [13] Brown, A. (1991). Functional load and the teaching of pronunciation. In A. Brown (Ed.), Teaching English pronunciation (pp. 221-224). Routledge. [14] Catford, J. C. (1987). Phonetics and the teaching of pronunciation: A systemic description of English phonology. In J. Morley (Ed.), Current perspectives on pronunciation (pp. 87-100). TESOL. . Pairs of sounds with a high functional load imply that failure to distinguish them is likely to induce communication issues. Therefore, among vowels and consonants, pairs with a high functional load should be prioritized for instruction and practice. Among vowels, /iː/−/ɪ/ (leave−live, feet−fit, reach−rich) is a pair with high functional load, whereas /uː/−/ʊ/ is a pair with low functional load. Regarding consonants, /l/−/r/ (light−right, cloud−crowd, play−pray) is a pair with a high functional load, whereas /θ/−/s/ is a pair with a low functional load.

  • Vowel pairs

    High functional load
    /iː/−/ɪ/ (leave−live, feet−fit, reach−rich)
    Low functional load
    /uː/−/ʊ/ (pool−pull, fool−full)
  • Consonant pairs

    High functional load
    /l/−/r/ (light−right, cloud−crowd, play−pray)
    Low functional load
    /θ/−/s/ (thick−sick, faith−face)
9

Contrastive analysis of English and Japanese sound systems

Comparison of English and Japanese sound systems

Contrastive analysis refers to the process of comparing a learner’s native language with the target language, focusing on their similarities and differences. When teaching English to learners whose native language is Japanese, having a precise understanding of the differences between Japanese and English in terms of vowel and consonant systems, syllable structure, rhythm, and intonation enables teachers to anticipate the challenges learners may experience and explicate the specific pronunciation issues that learners encounter.

10

Teaching vowels
(Differences between English and Japanese vowels)

English has more vowels than Japanese. For example, while Japanese has five vowels (/a, i, u, e, o/), English has more than 20 (though the number varies by accent). Consequently, native Japanese speakers need to learn to distinguish English vowel sounds that do not exist in Japanese. For example, the Japanese /i/ corresponds to the English /iː/ (leave) and /ɪ/ (live); the Japanese /e/ corresponds to the English /e/ (pen) and /eɪ/ (pain); the Japanese /u/ corresponds to the English /uː/ (pool) and /ʊ/ (pull); and the Japanese /o/ corresponds to English /ɔː/ (law) and /oʊ/ (low). Particular attention must be paid to the vowels corresponding to the Japanese /a/. There are many English vowels that correspond to Japanese /a/, including /æ/ (hat), /ʌ/ (hut), /ɑː/ (hot), /ɑɚ/ (heart), and /ɚː/ (hurt).

11

Teaching consonants
(Differences between English and Japanese consonants)

English has consonants that do not exist in Japanese, such as the fricatives /f/ (food), /v/ (voice), /θ/ (think), and /ð/ (the), as well as the approximants /l/ (light) and /r/ (right). Consequently, when Japanese speakers pronounce sounds in English, these consonants are often substituted with the closest Japanese consonants; /f/ tends to be replaced with [ɸ] (the “fu” sound), /v/ with /b/, /θ/ with /s/, and /ð/ with /(d)z/. Both English /l/ and /r/ are often replaced with the closest Japanese equivalent, an alveolar tap (/ɾ/), which can make it difficult to distinguish words such as light and right. Even for consonants common to both English and Japanese, there are some that are used differently depending on the language. For example, the nasal sound /ŋ/ used at the end of English words (king /kɪ́ŋ/) and the /s/ used before /i/ (sea /síː/) require practice for Japanese speakers.

12

Teaching spelling and pronunciation

In English, spelling and pronunciation do not have one-to-one correspondence. Many basic words (e.g., have, do, the) have irregular pronunciations; so one must memorize the pronunciation of each word as it is. However, there exist certain patterns in English spelling (cf. phonics). Regarding consonants, one must know the basic rules, such as <ph> = /f/ and <ch> = /tʃ/. Regarding vowels, memorizing rules such as <ai> = /eɪ/ (rain, aim, tail) and <au> = /ɔː/ (pause, August, audience) helps avoid pronunciation based on Romaji (Roman alphabet) reading.

Examples of consonant rules

<ph> = /f/ (phone, graph)
<ch> = /tʃ/ (children, teach)
<qu> = /kw/ (queen, quick)

Examples of vowel rules

<ai> = /eɪ/ (rain, aim, tail)
<au> = /ɔː/ (pause, August, audience)
<ou> = /aʊ/ (cloud, south, house)

13

Teaching connected speech

When words are used in natural speech, sounds may be connected (linking), change (assimilation), or disappear (elision). Japanese learners must pay special attention to this when listening to native English speakers. They should pay particular attention to the consonants /t/ and /d/. When a final /t/ is sandwiched between consonants, it may be elided (last minute /-s(t)m-/), and when /t/ is followed by /j/, coalescent assimilation may occur (last year /-tʃ-/). When /t/ is followed by a vowel, it is linked with the following vowel (sit up /-tʌp/). Additionally, when a plosive with the same place of articulation follows /t/, the release phase of a plosive is omitted (sit down /-td-/).

14

Teaching syllables and consonant clusters

There are significant differences between the syllable structures of English and Japanese. In English, consonants can occur consecutively before and after vowels, and there are not only open syllables (e.g., pie, tree) but also many closed syllables (e.g., desk, stop). In contrast, Japanese syllables generally comprise open syllables, either “consonant + vowel” or a single vowel, and there are no consonant clusters. Syllables ending in a consonant are extremely rare, such as those ending in /N/. Therefore, when native Japanese speakers produce English sounds in words, sentences and extended speech, they must avoid inserting a vowel after a consonant. For example, in English Christmas is pronounced as /krɪ́s.məs/ (= CCVC.CVC), comprising two syllables. However, when pronounced as kurisumasu in Japanese, it becomes /ku.ri.su.ma.su/ (= CV.CV.CV.CV.CV), comprising five syllables, with extra vowels inserted that do not exist in the English pronunciation.

15

Teaching word stress

In English, the position of stress is word-dependent, so learners must consult a dictionary to determine the correct stress placement. However, there are a set of rules, and the stress-suffix relationship is a particularly easy concept to teach. In English, there are: (i) suffixes that do not affect the position of stress, such as -ment or -ly (fórtunatefórtunately, commítcommítment); (ii) suffixes that carry the primary stress themselves, such as -eer and -ese (JapánJàpanése, éngineènginéer), and (iii) suffixes that shift the primary stress to the preceding syllable, such as -tion and -ic (éducàteèducátion, ecónomyèconómic).

  • (i) Suffixes that do not affect the position of stress

    <-ly> fórtunatefórtunately
    <-ment> commítcommítment

  • (ii) Suffixes that carry the primary stress themselves

    <-ese> JapánJàpanése
    <-eer> éngineènginéer

  • (iii) Suffixes that shift the primary stress to the preceding syllable

    <-tion> éducàteèducátion
    <-ic> ecónomyèconómic

16

Teaching rhythm

English stress-timed rhythm, in which stressed syllables occur at regular intervals, is characterized by a strong contrast between stressed and unstressed syllables. First, teachers must advise learners to use part-of-speech information to identify the content words that are pronounced strongly and the function words that are pronounced weakly in sentences. Next, they should teach them how to make stressed syllables strong, long, and clear. Additionally, they should teach them how to pronounce unstressed syllables, especially weak forms of function words, in a weak and quick manner. Japanese has a mora-timed rhythm. It is noteworthy that native Japanese speakers tend to pronounce all words with roughly the same intensity when speaking English.

17

Teaching intonation

Teaching intonation comprises three key points: (1) division of the sentences into tone units, (2) placement of focus, and (3) choice of tones. First, consider the sentence structure and have students think where to divide the sentences into tone units, for example, before a reading-aloud activity. When teaching placement of focus, begin by thinking which word should be emphasized based on the context. Additionally, it is necessary to ascertain whether students can both a) pronounce the sentence with the correct stresses and focus and b) identify which word the speaker is emphasizing within the sentence as listeners. Regarding intonation, teachers can teach rules such as the fact that wh-questions are often pronounced with a falling tone, or the rising tone can be employed in yes-no questions. In English, the pitch rises gradually from the stressed syllable of a focus word to the end of the sentence, whereas in Japanese interrogative sentences, the pitch rises only on the final particle.

  1. (1) Division of the sentences into tone units (1) Division of the sentences into tone units
  2. (2) Placement of focus (2) Placement of focus
  3. (3) Choice of tones (3) Choice of tones

References

  1. [1]

    Kachru, B. B. (1992). Teaching world Englishes. In B. B. Kachru (Ed.),The other tongue: English across cultures (2nd ed.) (pp. 355-365). Springer.

  2. [2]

    Derwing, T. M., & Munro, M. J. (1997). Accent, intelligibility, and comprehensibility: Evidence from four L1s. Studies in Second Language Acquisition, 19, 1–16.

  3. [3]

    Derwing, T. M. and M. J. Munro. (2015). Pronunciation fundamentals. John Benjamins Publishing Company.

  4. [4]

    Levis, J. M. (2005). Changing contexts and shifting paradigms in pronunciation teaching. TESOL Quarterly, 39(3), 369-377.

  5. [5]

    Levis, J. M. (2020). Revisiting the intelligibility and nativeness principles. Journal of Second Language Pronunciation, 6(3), 310-328.

  6. [6]

    杉本淳子・内田洋子 (2020).「英語教員養成における音声学教育:日本人英語教員のための〈教職音声学〉試案」『音声研究』24, 22-35.

  7. [7]

    Sugimoto, J., & Uchida, Y. (2018). Accentedness and acceptability ratings of Japanese English teachers’ pronunciation. Proceedings of PSLLT-9, 30-40.

  8. [8]

    Uchida, Y., & Sugimoto, J. (2020). Non-native English teachers’ confidence in their own pronunciation and attitudes towards teaching: A questionnaire survey in Japan. International Journal of Applied Linguistics, 30(1), 19-34.

  9. [9]

    Muller Levis, G., & Levis, J. (2016). Integrating pronunciation into listening/speaking classes. In T. Jones (Ed.) Pronunciation in the classroom: The overlooked essential (pp. 27-42). TESOL.

  10. [10]

    Uchida, Y., & Sugimoto, J. (2026). Developing pedagogical knowledge for English pronunciation teaching through material evaluation. Speak Out!, 74, 33-45.

  11. [11]

    Isbell, D. R., & Sakai, M. (2022). Pronunciation assessment in classroom contexts. In J. Levis, T. Derwing, & S. Sonsaat-Hegelheimer (Eds.), Second language pronunciation: Bridging the gap between research and teaching (pp. 194-214). John Wiley & Sons.

  12. [12]

    常本亜希 (2025). 「評価において何をどのように捉えるべきか」シンポジウム「教員のための音声指導と評価」外国語教育メディア学会 (LET) 第64回 (2025) 年次研究大会.

  13. [13]

    Brown, A. (1991). Functional load and the teaching of pronunciation. In A. Brown (Ed.), Teaching English pronunciation (pp. 221-224). Routledge.

  14. [14]

    Catford, J. C. (1987). Phonetics and the teaching of pronunciation: A systemic description of English phonology. In J. Morley (Ed.), Current perspectives on pronunciation (pp. 87-100). TESOL.