You can find the right rule and still lose marks because you misread a symbol. Linguistics olympiad data arrives in a mixture of IPA, practical romanisation and glossing abbreviations, and the problem’s own key overrides all of it. This guide covers what the marks above, below and beside a letter usually mean, which ones are most often dropped when students copy data, and which pinyin habits work against you.
Losing marks to the symbols, not to the language
The official IOL FAQ states that the individual contest contains “precisely five problems” and that “the time for the individual contest is normally six hours” (confirm current details on ioling.org). That is roughly seventy minutes per problem, all of it spent on data you have never seen, written in a notation nobody explained to you in school.
Here is the failure mode nobody warns students about. A solver reads a data set, forms a hypothesis, tests it, and finds it working on four items out of six. The two failures look like exceptions, so they invent a second rule to cover them. In fact the two “exceptions” were identical to the other four — the difference was a superscript mark they never registered. The analysis was destroyed by a reading error, not a reasoning error, and no amount of extra thinking recovers it because the thinking was never the problem.
In our own training sessions, three symbols get silently dropped more than any others when students copy data onto their working sheet: the superscript ʰ, the glottal stop ʔ, and the length mark ː. All three are small, all three are easy to read as decoration, and all three routinely carry the entire grammatical contrast a problem is built on. That is an editorial observation from coaching rather than a published statistic, but it matches what you would expect: the marks that look least like letters are the ones the eye skips.
So before you decode anything, work out which of four notation systems you are looking at. Problems mix them freely, sometimes within one data set.
- IPA (the International Phonetic Alphabet). One symbol, one sound, square brackets or slashes are common. This is the system the diacritic table below assumes.
- A practical romanisation. A writing system designed by missionaries, colonial administrators or the speech community itself. Here letters are reused freely: c might be [k], [tʃ] or [ts]; j might be the English j or the English y. You cannot import assumptions.
- Interlinear glosses. A second line under the data giving morpheme-by-morpheme meanings, usually with hyphens and small-capital abbreviations.
- The problem’s own key. A short paragraph, often in small print at the bottom, defining exactly the symbols the author thinks you might misread. It outranks everything else on this page.
Read the key first — out loud, if you are practising alone. It takes thirty seconds and it is the single highest-return half-minute in the paper. If you are still building your picture of how the contest is put together, start with our overview of what the International Linguistics Olympiad is and come back to this.
The diacritic survival table
These are default meanings, not guarantees. Traditions differ, and problem authors sometimes redefine a symbol deliberately to stop you coasting on prior knowledge.
| Symbol | Usually signals | Rough sound / example | The trap |
|---|---|---|---|
| ː | Length (long vowel or consonant) | aː is a held a | Counting aː as two vowels wrecks your syllable count |
| ʰ (superscript) | Aspiration | kʰ, like the k in English kit | If both k and kʰ appear, that contrast is almost certainly doing work |
| ʼ (after a consonant) | Ejective in many traditions; glottalisation or palatalisation in others | tʼ, kʼ | The symbol most often redefined by the key — never assume |
| ʔ | Glottal stop | The catch in uh-oh | Skipping it when copying changes the word into a different word |
| ŋ | Velar nasal | The ng in sing | It is one segment, not n + g |
| ɲ / ñ | Palatal nasal | Roughly the ny in canyon | Often patterns with other palatals, not with n |
| ◌̃ (tilde over a vowel) | Nasalisation | ã, õ | Frequently spreads from a neighbouring nasal — that spread is the rule |
| ◌́ / ◌̀ | High / low tone, or stress | á, à | Which one depends on the tradition; the key decides |
| ā (macron) | Long vowel in many romanisations; mid tone in some | ā, ī | The same mark carries different meanings in different problems |
| ◌̣ (dot below) | Retroflex in Indic romanisations; other uses elsewhere | ṭ, ḍ | Easy to lose entirely in handwriting |
| . (inside a word) | Syllable boundary | ka.ne | Not punctuation — it is data about structure |
| – (hyphen) | Morpheme boundary | ka-ne-tu | The author has already done part of the segmentation for you: use it |
| = | Clitic boundary | ka=ne | Signals an element that attaches but behaves differently from an affix |
| Ø | Zero — nothing where something could be | walk-Ø | A real part of the analysis, not a typo |

The pinyin trap: when Chinese literacy works against you
Every Chinese student arrives at a linguistics olympiad already fluent in one romanisation: Hanyu Pinyin. That is genuinely an advantage — you are used to the idea that Latin letters are a code for sounds rather than the sounds themselves, which many monolingual English readers find unnatural. But pinyin assigns several letters values that are unusual across the world’s romanisations, and those associations fire automatically when you read data.
| Letter | Pinyin habit says | Commonly transcribed in IPA as | What it is more likely to be in olympiad data |
|---|---|---|---|
| q | a ch-like sound | [tɕʰ] | [q], a uvular stop — a k made further back |
| x | a sh-like sound | [ɕ] | [x] or [ʃ] — a raspy h, or plain sh |
| c | [ts] | [tsʰ] | [k], [tʃ] or [s], depending entirely on the romanisation |
| j | a soft j | [tɕ] | [j] — the English y sound, extremely common worldwide |
| zh | an English j | [ʈʂ] | [ʒ] — the middle sound of measure |
| e | an uh vowel | [ɤ] | [e] — a plain mid front vowel |
| ü | a special u | [y] | [y] here too, but in some systems just a long or tense u |
| ' (apostrophe) | a syllable divider, as in Xi'an | — | An ejective, glottalisation or a glottal stop — a consonant, not punctuation |
The apostrophe row is the expensive one. A Chinese reader trained on Xi’an sees an apostrophe and reads “the syllables split here” — that is, nothing sounded. In most transcription traditions that mark is a segment: a consonant with its own distribution, its own rules, and often a leading role in the answer.
The fix is not to unlearn pinyin. It is to add one deliberate step: when you first meet a Latin letter in unfamiliar data, ask “what does this letter do in this problem?” rather than “what does this letter sound like?” You almost never need the phonetic value to state a correct rule.
Glosses, hyphens and abbreviation strings
Morphology and syntax problems often supply an interlinear gloss — a line under the data that maps each piece to a meaning. The widely used Leipzig conventions work like this:
- Hyphens mark morpheme boundaries, and the two lines must line up: three hyphen-separated pieces in the data means three hyphen-separated pieces in the gloss.
- Small capitals mark grammatical categories rather than dictionary meanings: PST (past), PL (plural), NOM (nominative), ERG (ergative), 1SG (first person singular).
- A period inside a gloss usually means one form carries two categories at once: 3SG.PST is a single ending expressing both.
- An equals sign marks a clitic — an element that leans on a neighbouring word without being a normal affix.
Two habits pay for themselves. First, count the pieces on both lines before reasoning; a mismatch means you have mis-segmented, and it is far cheaper to catch that in minute two than in minute forty. Second, treat the abbreviations as labels rather than theory. You do not need to know what an ergative is to notice that ERG appears on exactly the subjects of one class of verb — and noticing that is the answer.
A worked mini-drill: one misread symbol, one wrong rule
Here is an invented data set — not a past problem — in a language we will call Tavuri. Present tense on the left, past on the right.
| Present | Meaning | Past |
|---|---|---|
| kápu | he sees | kʰápu |
| tíra | he runs | tʰíra |
| píla | he sings | pʰíla |
| ámo | he sleeps | hámo |
Read carefully and the analysis takes a minute: the past tense aspirates the initial consonant, and when the word begins with a vowel and has no consonant to aspirate, an h is prefixed instead. One rule, one principled extension — a strong, general statement.
Now read carelessly. If the superscript ʰ never registers, the first three pairs look identical, and you are forced into a conclusion like “Tavuri does not mark past tense on these verbs” — with the fourth pair as an unexplained oddity. Everything you write afterwards is downstream of a reading failure, and no re-checking of your logic will find it, because your logic was fine.
This is why the very first pass over a data set should be mechanical rather than clever: circle every symbol that is not a plain letter, and write in the margin what you think each one is doing. Our past problems practice pack is the right place to run that drill, because the worked solutions we have for many of the years let you check whether the symbol you ignored was load-bearing.

Two weeks to notation fluency
Notation is one of the few parts of olympiad preparation that responds to short, boring, daily practice. A fortnight is enough to stop losing marks to it.
- Days 1–3: build your own symbol sheet. One side of A4, handwritten. Copying by hand is the point — it forces you to notice that ŋ has a descender and n does not.
- Every day: a ten-minute copying drill. Take ten transcribed words from any source, copy them by hand, then check character by character against the original. Score yourself on marks reproduced, not on speed.
- Every practice problem: a symbol pass first. Before hypothesising, circle every non-plain-ASCII character and annotate what you think it does. Thirty seconds, and it changes what you see.
- Handwriting discipline. The IOL FAQ’s kit list is pencils, sharpeners and erasers, so your answers are handwritten. If your ŋ can be read as an n, or your ʔ as a stray flick, you have created ambiguity in the one document a marker actually sees.
- Work with a partner. One person describes a transcription aloud, the other writes it; then swap and compare. This surfaces exactly the marks you habitually skip.
Fold these into whatever training cycle you are running — they sit comfortably inside the warm-up slot of our twelve-week study plan without displacing problem-solving time.
Frequently asked questions
Do I need to memorise the whole IPA chart?
No. You need the marks that recur — length, aspiration, nasalisation, tone, glottal stop — and the habit of reading the problem’s key before the data.
Can I bring an IPA chart or a dictionary into the contest?
The IOL FAQ lists pencils, sharpeners and erasers to bring, and bans phones, tablets and laptops. For anything else, confirm on ioling.org and with your national organiser.
Does knowing pinyin help or hurt?
Both. It builds comfort with sound-to-letter mapping, but several pinyin values are unusual worldwide, so check each letter against the problem rather than your instinct.
What if I still cannot identify a symbol during the contest?
Give it a neutral label, describe where it occurs, and state your rule in those terms. A correct generalisation does not require the phonetic name.
IOL Club is an independent guide operated by Hanlin Education for China-based international-school students. We are NOT affiliated with, endorsed by, or sponsored by the IOL Board. Competition format, eligibility, dates and rules change: confirm current details on ioling.org before acting on anything here. Errors reported to our editorial desk are corrected within 7 working days.
Keep reading.
From Three Hours to Six: What Actually Changes Between a National Round Paper and an IOL Paper
A national open round gives you roughly 25 minutes a problem; the IOL gives you over an hour. Four things that change…
How to Mark Your Own Linguistics Olympiad Practice Paper When There Is No Solution Sheet
Worked solutions do not exist for every past paper. A five-band self-marking rubric, the regeneration test, and the four ways students quietly…
The Deadline Before the Deadline: Planning a 2026-27 Linguistics Olympiad Autumn Backwards From the Sitting
Registration for the coming cycle is reported to close earlier than usual. An autumn planned backwards from the sitting, with compression options…

