On this page

You know the words because you met them on a page: spelled out, spaced apart, with all the time you needed to look. Spoken Japanese gives you none of that. The gaps between words are gone, the vowels are half-swallowed, and whole phrases arrive fused into a single sound, so the sentence is over before you have found where one word ends and the next begins. The vocabulary is not the problem. The skill you never trained is the one that turns a run of sound back into the words you already know.

You learned the words to be seen, not heard

The way you first meet a word decides how you can get it back. A flashcard, a subtitle, a dictionary entry: every one of them is the word standing still, in isolation, with its spelling in plain view. Recognising it there is one skill. Catching it as it flies past inside a sentence, with no spelling and no pause, is a completely different one, and being good at the first does not hand you the second. That is why a word you would name instantly on a card can go by three times in a conversation without registering. You have the word. You do not yet have it as a sound in motion.

Reading also lets you set your own pace. Your eye can stop, back up, sit on a hard word and return to the sentence. Listening runs at the speaker's pace, forward only, and never waits. So even when every word in a sentence is one you know, you can still lose the sentence, simply because you were still decoding word two when word five arrived.

Japanese speech has no gaps, and your ear is hunting for them

Speech in any language is a continuous stream. The little silences you think you hear between words are mostly not there. You insert them, because you already know where the boundaries fall. Linguists call this segmentation, and it is the first thing to break when you listen to a language you learned mainly by reading.

In writing, Japanese hands you a boundary cue for free. The switch between kanji and kana shows where content words stop and grammar begins: in 私は学生です, the kanji 私 and 学生 are the words, and the kana は and です are the joints between them. Your eye uses that switching to carve the sentence up without any effort. In speech the cue is gone. わたしわがくせいです arrives as one unbroken ribbon of sound, and nothing in it announces where 私 ends and は begins. You have to place every boundary yourself, from memory, in real time.

Left panel: the sentence written as 私は学生です, where the kanji-and-kana switching marks each word. Right panel: the same sentence written as one unbroken kana stream, わたしわがくせいです, with no visible boundaries.
In writing the seams are visible; in speech you supply them yourself.

English speakers meet a second obstacle here. In English you find word boundaries partly through stress, since a strong syllable tends to start a new word. Japanese does not work that way. It is timed in even beats, or morae, and its main boundary music is pitch rather than stress. If you never learned to hear pitch, you are trying to segment Japanese with an English tool the language does not answer to.

Natives don't say the word you memorised

Even once you find the boundary, the word inside it is often not the one on your flashcard. Fast, casual speech wears words down. Vowels go quiet: です lands closer to "des", 好き closer to "ski", the u barely voiced at all. Whole grammatical endings collapse. ている becomes てる, then てん before の. ておく becomes とく. なければ becomes なきゃ. てしまう becomes ちゃう. The dictionary form you studied is frequently the least likely shape you will actually hear it in.

A table comparing textbook forms with what you hear: 食べている becomes 食べてる, 行かなければ becomes 行かなきゃ, 忘れてしまう becomes 忘れちゃう, 買っておく becomes 買っとく, 何しているの becomes 何してんの.
None of this is slang; it is ordinary speech at ordinary speed.

None of these are slang you are free to skip. They are how ordinary sentences sound at ordinary speed, from a newsreader as much as from a friend. If you only ever trained on the full, careful citation form, then every contraction is a word your ear has genuinely never met, even though you know it cold in writing. Loanwords get the same treatment and worse, because they were already reshaped once on the way into the language, which is a separate rabbit hole entirely.

Fusion happens above the word level too, and it is easier to miss because nothing there looks like slang. 食べ過ぎる and 話し始めた are two full verbs welded into one, and a listener hears the result as a single unit, not as a known verb followed by a second one you could pause and look up.

The unit that arrives is a chunk, not a word

Natives do not speak word by word, and they do not listen word by word either. They handle speech in phrase-sized chunks, taking a whole grammatical unit in at once. A beginner does the reverse: decode one word, hold it, decode the next, try to keep the first in mind. At conversational speed that queue overflows in about a second, and the sentence dissolves into noise. Not because it was too fast for your ears, but because it was too fast for your decoder.

The phrase 何してんの broken into its parts: 何 (what), し (the verb する), て (linking te), ん (worn down from いる), の (makes it a question).
One phrase, five jobs, three sounds — and no gap anywhere to grab.

This is why slowing the audio down helps less than you would hope. Half-speed makes each sound clearer, but it does not build the chunks. What builds them is hearing the same collocations and the same sentence frames so often that a run of five words becomes one familiar shape you recognise whole. Until then you are spelling out every sentence, in a manner of speaking, while the speaker reads fluently.

Pitch and rhythm are the cues you skipped

Here is the cue most learners never install. In connected Japanese, pitch does more than separate 箸 from 橋; it also helps mark where words begin, because each word tends to carry its own small pitch shape. Rhythm does related work: the even, mora-timed beat is a grid your ear can lock onto, and the quiet, devoiced vowels fall into predictable slots within it. Native listeners lean on all of this without noticing. If pitch is something you have only ever read about, you are missing a whole channel the language uses to hand you its boundaries. It is worth knowing when pitch accent is worth learning and when it is not, because for listening specifically the answer comes earlier than most people assume.

What to train, and how

The fix is not more vocabulary. It is exposure to the sound of the words you already have, at real speed, until your ear stops needing the page. Three things move it.

Listen to speech made for natives, not for learners. Textbook audio is over-articulated on purpose, so it trains you to understand a register nobody actually speaks. A drama, a podcast, an interview, even a minute at a time, puts the missing gaps, the reductions and the rhythm in front of you all at once.

Listen to things you already understand on the page. Read a passage first, then hear it. Now your ear can spend its whole effort on the sound instead of the meaning, and the two versions of the word, the seen one and the heard one, finally get stored together instead of apart.

Turn recognition into recall. Playing a clip until it clicks is comprehension; being able to catch the same word cold next week is memory, and memory only comes from retrieval, from meeting the word again on a schedule rather than replaying it once and moving on.

This is where how you keep a word starts to matter. A word stored as spelling and meaning alone is a word you can read. A word stored with its sound attached is a word you can catch. It is why the cards Kikusho builds carry native audio and example sentences you can play: when you review, you hear the word rather than only seeing it, so the ear collects the same repetitions the eye does. And because you can capture by voice, saying the phrase you just half-caught, the words that land in your deck are the exact ones your ear keeps failing on, instead of a generic list someone else drew up.

Common questions

Why can I read Japanese but not understand it spoken?

Because reading and listening are different skills. Reading gives you spelling, spaces you supply from knowledge, and time to stop; listening gives you a continuous stream at the speaker's pace, with the vowels worn down and the endings contracted. Knowing a word on the page does not mean your ear can catch it going past.

Does slowing the audio down help?

A little. Half-speed audio makes individual sounds clearer, but it does not build the phrase-sized chunks that let you take a sentence in at once. Repeated exposure to natural speech does that; slowed audio mainly buys you time to decode word by word, which is the habit you are trying to grow out of.

Should I learn pitch accent to understand speech better?

For listening it helps earlier than most people think, because pitch is one of the cues that marks where words begin in a continuous stream. You do not need to drill every pattern, but training your ear to hear pitch at all gives you a boundary signal Japanese uses and English does not.

What should I actually listen to?

Speech made for natives rather than for learners: dramas, podcasts, interviews, in short doses. Textbook audio is over-articulated on purpose and trains you for a register nobody speaks. Listening to something you have already read is especially efficient, because your ear can spend its effort on the sound instead of the meaning.