News IOL Club · Bulletin

How to Read Linguistics Olympiad Data: IPA, Diacritics and the Pinyin Trap (2026)

5 Aug 2026 · 10 min read · By the IOL Club editorial team

You can find the right rule and still lose marks because you misread a symbol. Linguistics olympiad data arrives in a mixture of IPA, practical romanisation and glossing abbreviations, and the problem’s own key overrides all of it. This guide covers what the marks above, below and beside a letter usually mean, which ones are most often dropped when students copy data, and which pinyin habits work against you.

Losing marks to the symbols, not to the language

The official IOL FAQ states that the individual contest contains “precisely five problems” and that “the time for the individual contest is normally six hours” (confirm current details on ioling.org). That is roughly seventy minutes per problem, all of it spent on data you have never seen, written in a notation nobody explained to you in school.

Here is the failure mode nobody warns students about. A solver reads a data set, forms a hypothesis, tests it, and finds it working on four items out of six. The two failures look like exceptions, so they invent a second rule to cover them. In fact the two “exceptions” were identical to the other four — the difference was a superscript mark they never registered. The analysis was destroyed by a reading error, not a reasoning error, and no amount of extra thinking recovers it because the thinking was never the problem.

In our own training sessions, three symbols get silently dropped more than any others when students copy data onto their working sheet: the superscript ʰ, the glottal stop ʔ, and the length mark ː. All three are small, all three are easy to read as decoration, and all three routinely carry the entire grammatical contrast a problem is built on. That is an editorial observation from coaching rather than a published statistic, but it matches what you would expect: the marks that look least like letters are the ones the eye skips.

So before you decode anything, work out which of four notation systems you are looking at. Problems mix them freely, sometimes within one data set.

  • IPA (the International Phonetic Alphabet). One symbol, one sound, square brackets or slashes are common. This is the system the diacritic table below assumes.
  • A practical romanisation. A writing system designed by missionaries, colonial administrators or the speech community itself. Here letters are reused freely: c might be [k], [tʃ] or [ts]; j might be the English j or the English y. You cannot import assumptions.
  • Interlinear glosses. A second line under the data giving morpheme-by-morpheme meanings, usually with hyphens and small-capital abbreviations.
  • The problem’s own key. A short paragraph, often in small print at the bottom, defining exactly the symbols the author thinks you might misread. It outranks everything else on this page.

Read the key first — out loud, if you are practising alone. It takes thirty seconds and it is the single highest-return half-minute in the paper. If you are still building your picture of how the contest is put together, start with our overview of what the International Linguistics Olympiad is and come back to this.

The diacritic survival table

These are default meanings, not guarantees. Traditions differ, and problem authors sometimes redefine a symbol deliberately to stop you coasting on prior knowledge.

Symbol Usually signals Rough sound / example The trap
ː Length (long vowel or consonant) aː is a held a Counting as two vowels wrecks your syllable count
ʰ (superscript) Aspiration kʰ, like the k in English kit If both k and appear, that contrast is almost certainly doing work
ʼ (after a consonant) Ejective in many traditions; glottalisation or palatalisation in others tʼ, kʼ The symbol most often redefined by the key — never assume
ʔ Glottal stop The catch in uh-oh Skipping it when copying changes the word into a different word
ŋ Velar nasal The ng in sing It is one segment, not n + g
ɲ / ñ Palatal nasal Roughly the ny in canyon Often patterns with other palatals, not with n
◌̃ (tilde over a vowel) Nasalisation ã, õ Frequently spreads from a neighbouring nasal — that spread is the rule
◌́ / ◌̀ High / low tone, or stress á, à Which one depends on the tradition; the key decides
ā (macron) Long vowel in many romanisations; mid tone in some ā, ī The same mark carries different meanings in different problems
◌̣ (dot below) Retroflex in Indic romanisations; other uses elsewhere ṭ, ḍ Easy to lose entirely in handwriting
. (inside a word) Syllable boundary ka.ne Not punctuation — it is data about structure
(hyphen) Morpheme boundary ka-ne-tu The author has already done part of the segmentation for you: use it
= Clitic boundary ka=ne Signals an element that attaches but behaves differently from an affix
Ø Zero — nothing where something could be walk-Ø A real part of the analysis, not a typo
Default readings of common transcription marks. Where a problem supplies its own key, the key always wins.
Diagram of one invented transcribed word, with six labelled cards explaining the superscript h, acute accent, length mark, syllable period, velar nasal and glottal stop
Anatomy of a transcribed word — invented example, built for this guide.

The pinyin trap: when Chinese literacy works against you

Every Chinese student arrives at a linguistics olympiad already fluent in one romanisation: Hanyu Pinyin. That is genuinely an advantage — you are used to the idea that Latin letters are a code for sounds rather than the sounds themselves, which many monolingual English readers find unnatural. But pinyin assigns several letters values that are unusual across the world’s romanisations, and those associations fire automatically when you read data.

Letter Pinyin habit says Commonly transcribed in IPA as What it is more likely to be in olympiad data
q a ch-like sound [tɕʰ] [q], a uvular stop — a k made further back
x a sh-like sound [ɕ] [x] or [ʃ] — a raspy h, or plain sh
c [ts] [tsʰ] [k], [tʃ] or [s], depending entirely on the romanisation
j a soft j [tɕ] [j] — the English y sound, extremely common worldwide
zh an English j [ʈʂ] [ʒ] — the middle sound of measure
e an uh vowel [ɤ] [e] — a plain mid front vowel
ü a special u [y] [y] here too, but in some systems just a long or tense u
' (apostrophe) a syllable divider, as in Xi'an An ejective, glottalisation or a glottal stop — a consonant, not punctuation
Broad transcriptions of pinyin values, alongside the readings the same letters more often carry elsewhere. Conventions vary; the problem’s key decides.

The apostrophe row is the expensive one. A Chinese reader trained on Xi’an sees an apostrophe and reads “the syllables split here” — that is, nothing sounded. In most transcription traditions that mark is a segment: a consonant with its own distribution, its own rules, and often a leading role in the answer.

The fix is not to unlearn pinyin. It is to add one deliberate step: when you first meet a Latin letter in unfamiliar data, ask “what does this letter do in this problem?” rather than “what does this letter sound like?” You almost never need the phonetic value to state a correct rule.

Scan · IOL Club mentor

Not sure where to start?

One WhatsApp message to the club mentor — ask about:

  • Registration help for your national olympiad
  • Free past-paper archive (22 years) + solutions
  • Weekly walkthroughs + a mentor for your grade
Scan the WeChat QR code
WhatsApp
WhatsApp
WeChat
WeChat

Name + school year + country

Glosses, hyphens and abbreviation strings

Morphology and syntax problems often supply an interlinear gloss — a line under the data that maps each piece to a meaning. The widely used Leipzig conventions work like this:

  • Hyphens mark morpheme boundaries, and the two lines must line up: three hyphen-separated pieces in the data means three hyphen-separated pieces in the gloss.
  • Small capitals mark grammatical categories rather than dictionary meanings: PST (past), PL (plural), NOM (nominative), ERG (ergative), 1SG (first person singular).
  • A period inside a gloss usually means one form carries two categories at once: 3SG.PST is a single ending expressing both.
  • An equals sign marks a clitic — an element that leans on a neighbouring word without being a normal affix.

Two habits pay for themselves. First, count the pieces on both lines before reasoning; a mismatch means you have mis-segmented, and it is far cheaper to catch that in minute two than in minute forty. Second, treat the abbreviations as labels rather than theory. You do not need to know what an ergative is to notice that ERG appears on exactly the subjects of one class of verb — and noticing that is the answer.

A worked mini-drill: one misread symbol, one wrong rule

Here is an invented data set — not a past problem — in a language we will call Tavuri. Present tense on the left, past on the right.

Present Meaning Past
kápu he sees kʰápu
tíra he runs tʰíra
píla he sings pʰíla
ámo he sleeps hámo
Invented illustrative data. The whole tense contrast lives in a superscript mark and one prefixed letter.

Read carefully and the analysis takes a minute: the past tense aspirates the initial consonant, and when the word begins with a vowel and has no consonant to aspirate, an h is prefixed instead. One rule, one principled extension — a strong, general statement.

Now read carelessly. If the superscript ʰ never registers, the first three pairs look identical, and you are forced into a conclusion like “Tavuri does not mark past tense on these verbs” — with the fourth pair as an unexplained oddity. Everything you write afterwards is downstream of a reading failure, and no re-checking of your logic will find it, because your logic was fine.

This is why the very first pass over a data set should be mechanical rather than clever: circle every symbol that is not a plain letter, and write in the margin what you think each one is doing. Our past problems practice pack is the right place to run that drill, because the worked solutions we have for many of the years let you check whether the symbol you ignored was load-bearing.

Four-step flow for handling an unfamiliar transcription symbol: check the key, check the contrast, name it neutrally, copy it exactly
A four-step routine for unfamiliar notation, usable on any language and any problem type.

Two weeks to notation fluency

Notation is one of the few parts of olympiad preparation that responds to short, boring, daily practice. A fortnight is enough to stop losing marks to it.

  • Days 1–3: build your own symbol sheet. One side of A4, handwritten. Copying by hand is the point — it forces you to notice that ŋ has a descender and n does not.
  • Every day: a ten-minute copying drill. Take ten transcribed words from any source, copy them by hand, then check character by character against the original. Score yourself on marks reproduced, not on speed.
  • Every practice problem: a symbol pass first. Before hypothesising, circle every non-plain-ASCII character and annotate what you think it does. Thirty seconds, and it changes what you see.
  • Handwriting discipline. The IOL FAQ’s kit list is pencils, sharpeners and erasers, so your answers are handwritten. If your ŋ can be read as an n, or your ʔ as a stray flick, you have created ambiguity in the one document a marker actually sees.
  • Work with a partner. One person describes a transcription aloud, the other writes it; then swap and compare. This surfaces exactly the marks you habitually skip.

Fold these into whatever training cycle you are running — they sit comfortably inside the warm-up slot of our twelve-week study plan without displacing problem-solving time.

Frequently asked questions

Do I need to memorise the whole IPA chart?
No. You need the marks that recur — length, aspiration, nasalisation, tone, glottal stop — and the habit of reading the problem’s key before the data.

Can I bring an IPA chart or a dictionary into the contest?
The IOL FAQ lists pencils, sharpeners and erasers to bring, and bans phones, tablets and laptops. For anything else, confirm on ioling.org and with your national organiser.

Does knowing pinyin help or hurt?
Both. It builds comfort with sound-to-letter mapping, but several pinyin values are unusual worldwide, so check each letter against the problem rather than your instinct.

What if I still cannot identify a symbol during the contest?
Give it a neutral label, describe where it occurs, and state your rule in those terms. A correct generalisation does not require the phonetic name.

IOL Club is an independent guide operated by Hanlin Education for China-based international-school students. We are NOT affiliated with, endorsed by, or sponsored by the IOL Board. Competition format, eligibility, dates and rules change: confirm current details on ioling.org before acting on anything here. Errors reported to our editorial desk are corrected within 7 working days.

Scan · IOL Club mentor

Not sure where to start?

One WhatsApp message to the club mentor — ask about:

  • Registration help for your national olympiad
  • Free past-paper archive (22 years) + solutions
  • Weekly walkthroughs + a mentor for your grade
Scan the WeChat QR code
WhatsApp
WhatsApp
WeChat
WeChat

Name + school year + country

More from the club

Keep reading.

All bulletins →