On this page
You tell two look-alike kanji apart by meeting each one inside a word it actually appears in, because look-alike characters almost never compete for the same slot in a real word. 待 and 持 are genuine twins standing alone and no trouble at all in 待つ and 持つ: by the time you reach the つ, the sentence has already settled whether someone is waiting or holding.
Which means the confusion is not a property of the characters. It is a property of the card that showed you one of them alone, on a blank field, with nothing around it to decide the question.
Look-alike kanji look alike because they are related
Most confusable pairs are not coincidental twins. They are relatives sharing a component, and the component they share is usually the one that once carried the sound.
待, 持 and 時 all contain 寺. That is not an accident of drawing. 寺 was the phonetic element, the part that told a reader roughly how the character was pronounced. What differs is the piece on the left — 彳, the radical for stepping or going; 扌, the hand; 日, the sun — and that piece is the one carrying the meaning.

So the difference between two look-alikes is never cosmetic. It is the entire semantic content of the character. 持 has a hand in it because holding is something you do with hands. 待 has a road in it because waiting was standing about on one. The stroke you keep skipping over is the only part that says what the character is about.
The sound half has eroded, and it is worth knowing that it has. 持 is ジ and 時 is ジ, but 待 is タイ and 特 is トク. The phonetic component is a fossil from a much older stage of Chinese, not a rule you can lean on — which is one more reason a reading only resolves inside the word it belongs to. What survives intact is the other half. The radical still tells you the domain, reliably, across thousands of characters.
The pair only exists in the place you built it
Here is what the confusable-pairs lists leave out. Go looking for a real word in which 待 and 持 could both plausibly stand. There isn't one.

Read down the middle column, then the right-hand one. Nothing in either would take the other's character — not because a rule forbids it, but because the two characters mean unrelated things, and words are made of meanings. 休む is resting and 体 is a body. 未来 is what has not arrived yet; 週末 is the end of a week.
The pair is a pair in exactly one place: a study session where both were stripped of everything around them and set side by side. That is the single context in the entire language where they compete, and you built it yourself.
Studying the pair side by side is what welds it together
Nearly every article on this subject is a list of confusable pairs. The lists are accurate. They are also, mechanically, the thing that makes the pairs hard.
When you study 未 and 末 together, you are training a retrieval cue — the short-topped one with two horizontals — that points at two answers. Memory retrieves by cue. A cue resolving to one item is a memory; a cue resolving to two is a coin toss you will perform every single time, and each toss strengthens both candidates equally, because both were present when you made it.
This is why a pair often feels worse after the list than before it. Before, you knew 週末 perfectly well and had never once thought about 未. After, you have two characters filed under one shape and a gate between you and either of them.
The fix is not a cleverer mnemonic for the gate. It is to stop building the gate: give each character a different cue by giving it a different word.
A word hands you three cues the bare character withholds
Reading is not a shape-identification task. By the time your eye reaches a kanji, three separate things have usually already narrowed it to one candidate.
The kana trailing behind it. 大きい can only be 大, because 太 takes い on its own in 太い, and 犬 takes nothing at all, being a noun. The okurigana settles the character before you have finished looking at it.
The character beside it. In a compound, its neighbour has already fixed the domain. 土曜日 is a day of the week and 弁護士 is a profession, so the question of whether that glyph had a long top stroke or a short one never comes up. N3's abstract nouns lean on this even harder: 価値, 内容 and the rest are two kanji welded together, and reading the compound is reading what each half already means, not recognizing a new shape whole.
The slot in the sentence. Particles, position and what the sentence is doing rule out most of the field before shape is consulted at all.

That is the normal condition of reading, and it is why fluent readers do not experience these pairs as pairs. They are not resolving an ambiguity quickly. There is no ambiguity to resolve.
The pairs where the difference really is one stroke
Some pairs deserve the reputation. 土 and 士 differ only in which horizontal stroke is longer. 未 and 末 differ the same way. 千 and 干, 王 and 玉, 犬 and 大 come down to a dot or its absence. In these there is no radical to reason from, and no meaning hiding in the difference.
These are the ones worth a deliberate hook. But notice where the hook has to live: not on the character, on a word. 未 means not yet, and 未来 is the future, the thing that has not yet come. 末 is the tip or end of something, and 週末 is the end of the week. The hook works because a word gave it something to be about — on the bare characters, "the long stroke is on top" is a fact about ink with nothing to attach to.
Even here the word does most of the lifting on its own. You will meet 週末 constantly and 未 mainly in a small, coherent family: 未来, 未定, 未満. Meeting each one thirty times in its own company is what eventually makes the stroke length redundant, which is the point at which you stop needing the hook at all.
You are almost never asked to draw one from memory
There is a real version of this problem, and it is worth naming so you can tell it apart from the one you probably have.
Producing a character from nothing — writing 待 on paper with no prompt — genuinely requires you to hold the shape precisely. Recognising 待 inside 待っている requires almost nothing of the kind. These are different tasks with different difficulty, and the two decay at different rates — the second is the one you face every time you read a sign, a menu, a message or a page of manga.
Typing does not require the shape either. Type まつ and the IME offers you 待つ and 松, already formed, with their meanings implied by the words themselves. What it asks is whether you can recognise the right one when you see it, which is the recognition task again.
So if you are failing to distinguish two characters while reading, the problem is not that your visual memory is imprecise. It is that you learned them somewhere the context had been removed, and now you are trying to read the way you studied.
What to do the next time two of them collide
When a look-alike stops you mid-sentence, the instinct is to go and study the pair — pull up both characters, compare them, drill the difference. That instinct rebuilds the exact conditions that created the problem.
Do the opposite. Take the word you were actually reading when you stalled, not the character, and keep that. If you stumbled on 期待, keep 期待. The next time 待 appears in 待ち合わせ, keep that too. Two words, two contexts, two separate cues, and 持 was never in the room for either of them. This is the same move that works for readings and for particles, and the same reason words rather than characters are what a page actually asks of you.
That only works if keeping a word is cheap. If capturing 期待 means opening a dictionary, checking the reading, writing a card and filing it, you will decide it wasn't important and read on — and the word you actually met, in the sentence that actually confused you, is gone.
This is the part Kikusho exists to remove. Point the camera at the line, say it, or paste it, and a card comes back with the word, its reading, the meaning, a register note and an example sentence, which you check before it is saved. Any word inside that example sentence can be selected to become a card of its own, so a character reaches you as several words rather than as one glyph seen repeatedly. On a photo — a page, a sign, a menu — you can highlight up to five things and get five cards from the one capture. Reviews run on FSRS and work with no signal at all.
None of this is really about kanji. It is about the difference between a character and a word, and about which of the two your study is made of. A character seen in isolation is a shape with several possible meanings and no way to choose. A word is a shape with everything else attached — the reading, the sense, the grammar, the situation you met it in — and none of the attachments fit the look-alike.
Put the character back in a word and the pair has nowhere left to live.
Common questions
What is the fastest way to stop confusing two similar kanji?
Stop studying them together. Take each character and learn it inside a different everyday word — 休 in 休む, 体 in 体 — so that each one gets its own retrieval cue instead of sharing one shape between two answers. Two characters studied side by side are trained as a single ambiguous cue, which is why a pair often feels harder after you have compared them than before.
Why do so many kanji look almost identical?
Because they are usually relatives rather than coincidental twins. Many look-alikes share a phonetic component that once indicated the pronunciation — 待, 持 and 時 all contain 寺 — and differ in the radical, which is the part carrying the meaning. The stroke that seems like a minor detail is the entire semantic content of the character.
Are radicals worth learning to tell similar characters apart?
The common ones are, because they are what actually differs between look-alikes. Knowing that 扌 is the hand and 彳 is going or stepping makes 持 and 待 two unrelated ideas rather than two versions of one shape. The old phonetic components are much less useful, since the sounds have drifted: 持 is read ジ but 待 is タイ.
Do I need to tell similar kanji apart when they are on their own?
Almost never, and that matters because it is the hardest version of the task. Reading gives you the trailing kana, the neighbouring character and the grammatical slot, all of which narrow the field before shape is consulted. Typing does the same: an IME offers you formed candidates to recognise rather than asking you to produce the shape.



