News IOL Club · Bulletin

Set Your Own Linguistics Olympiad Problem: The Practice Method That Finds the Holes in Your Understanding (2026)

27 Aug 2026 · 9 min read · By the IOL Club editorial team

Solving a linguistics olympiad problem tests whether you can find a system someone else guaranteed was there. Building one removes the guarantee. You have to invent a system, generate data that determines it uniquely, and prove to yourself that nothing else fits. Students who try this once usually discover that several rules they thought they understood were only half-understood.

Why construction is harder than solution — and teaches more

When you solve, the setter has already done the hard part. They chose which forms to show you, in what order, and they checked that the data forces exactly one answer. You are handed a problem that is guaranteed to be solvable, and your job is search.

When you construct, all of that becomes yours. And the moment you try it, an uncomfortable thing happens: you find you cannot state your own rule precisely enough to generate data from it. That gap — between recognising a pattern and being able to run it forwards — is exactly the gap that costs marks under exam conditions, because the IOL individual contest asks you to write solutions, not just tick answers. The regulations set five problems in six hours, so roughly seventy minutes per problem, most of which is spent making a half-formed intuition explicit enough to defend on paper.

Construction attacks that gap directly. It is also cheap: no coach, no answer key, no internet, twenty minutes and a sheet of paper. And it produces something you can hand to a classmate, which makes it the single most useful thing to bring to a school club session.

Six-step construction loop: invent a small system, generate forms, strip the glosses, solve it yourself later, find the gaps, then repeat
The loop only pays off if you complete it. Steps 1 to 3 feel productive; steps 4 to 6 are where the misunderstandings surface.

The four constraints every real problem satisfies

Read enough competition problems and you notice they all obey the same quiet contract. Once you are setting rather than solving, that contract becomes a checklist.

Constraint What it means How a setter checks it What breaks if you skip it
Self-contained Everything needed is on the page; no outside knowledge of the language is required Give it to someone who has never met the language and forbid lookups The problem tests background, not reasoning — the one thing the format refuses to do
Determined Exactly one analysis fits every row of the data The delete-a-row test, then the rival-analysis check Two solvers give different correct answers and neither can be marked
Controlled ambiguity Residual ambiguity is allowed, but never in anything you ask about List what the data does not settle, then check no question touches it You penalise a solver for a gap you created
Graded Early rows give the skeleton; the discriminating rows come after Read your own data in order and mark where each rule becomes deducible The problem is either impossible at row one or trivial by row three
Sized A competent solver finishes in about an hour Time yourself in step 4 and compare Unusable for practice — the contest allows about seventy minutes per problem

The seventy-minute figure is arithmetic on the published format — five problems, six hours — not an official per-problem allowance. Treat it as a sanity check on scale, not a rule.

Scan · IOL Club mentor

Not sure where to start?

One WhatsApp message to the club mentor — ask about:

  • Registration help for your national olympiad
  • Free past-paper archive (22 years) + solutions
  • Weekly walkthroughs + a mentor for your grade
Scan the WeChat QR code
WhatsApp
WhatsApp
WeChat
WeChat

Name + school year + country

Build one in twenty minutes: a worked example

Here is a complete small system, invented for this article. Nothing below is drawn from any competition paper — make your own the same way rather than reusing this one.

Step 1 — invent the system. Write the rules out in full sentences, because half-sentences hide errors:

  • A word is root + (plural) + (locative), in that order. Either suffix may be absent.
  • Plural is -ter or -tar; locative is -sen or -san.
  • Front vowels are e, i; back vowels are a, o, u. A suffix takes its front form if the last vowel of the root is front, and its back form otherwise.

Step 2 — generate forms. Pick roots deliberately, not randomly. Two with uniform vowels, two with mixed vowels pointing in opposite directions:

Form Gloss (hidden from the solver) Job it does in the problem
kemi / kemiter / kemisen / kemitersen house / houses / in the house / in the houses Establishes the slots and their order
tolu / tolutar / tolusan river / rivers / in the river Reveals that suffixes come in two shapes
arni / arniter bird / birds Back vowel first, front vowel last — takes the front suffix
iklo / iklosan stone / in the stone Front vowel first, back vowel last — takes the back suffix

Step 3 — strip and shuffle. Remove the third column, scramble the row order, and ask questions that require the rule rather than recall: translate a form you never showed, and translate into the language a form whose suffix shape the solver must predict.

The delete-a-row test: is the data actually sufficient?

Now the part that teaches. Go through your data one row at a time and ask: if I removed this row, would any rule become undecidable? Rows that answer no are decoration — and rows that are the sole evidence for a rule are the load-bearing walls of your problem.

Grid showing which data rows are the sole evidence for each rule, revealing that the harmony rule depends on only two rows
Run this grid on your own data. Rules with a single supporting row are fragile; rules with none are not in your problem at all, however clearly they sit in your head.

Applied to the example, the grid says something useful. kemi and tolu have uniform vowels, so on their own they are compatible with three different rules: harmony keyed on the first vowel, on the last vowel, or on the whole root. Only arni (back vowel first, front last) and iklo (front first, back last) break the tie — and because they point in opposite directions, together they prove the rule rather than merely suggesting it. One discriminating datum is a claim; two pulling opposite ways is a proof.

Now the second test, and the one that catches setters out. Rival analysis: having found the intended answer, deliberately try to construct a different one that also fits. Do this on the example and you find a genuine hole. In kemitersen — the only form in the set carrying both suffixes — the locative appears in its front shape, and both the root’s last vowel (i) and the vowel immediately before the suffix (e) are front. The two candidate triggers agree, so nothing in the data separates them. The set is therefore equally consistent with two rules: harmony controlled by the root, or harmony spreading locally from whatever vowel precedes the suffix. The problem as drafted cannot tell them apart, and a solver who wrote either would be right.

That is a real defect, and the fix is instructive. Add a suffix that does not harmonise — say an invariant marker -mo — and then show a form where it sits between the root and the locative. If the language gives kemimosen, harmony is root-controlled; if it gives kemimosan, it spreads locally. One extra row, and the ambiguity is gone.

Notice what just happened. To repair your own problem you had to articulate a distinction — root-controlled versus local spreading — that you probably would not have thought about while solving. That is the whole return on the exercise.

Reading published papers, and a four-week rotation

Use published problems as a template, not a source. Study them for their architecture, never for their content. Copying somebody else’s data is both pointless as practice and not yours to reuse. What is genuinely worth stealing is structural, and you can read it off any paper in ten minutes:

  • How many rows before the first rule is deducible? Count them. You will usually find the answer is small — and that your own drafts are far too generous.
  • Where do the discriminating rows sit? Rarely at the top, rarely at the very bottom.
  • How many rules interact? Two interacting rules is a solid problem; four is usually a mess.
  • What do the questions ask for? Translation both ways is the standard pair, because it forces the solver to run the rule forwards as well as backwards.
  • What is deliberately left undetermined? Then check that nothing is asked about it.

For a supply of complete papers to read this way, we keep our own compiled pack of past problems, with worked solutions for some years rather than every year — scan the QR in the site footer to request it. If you are new to the problem families themselves, our overview of the competition is the place to start before you try to set anything.

Then rotate setting against solving. Setting is a supplement, not a replacement. Solving still has to dominate, because that is the actual skill being tested. Our recommended ratio is roughly four hours of solving to one hour of setting — enough to get the diagnostic benefit without displacing the main work.

  • Week 1 — build. Twenty minutes inventing a system in a family you find hardest, thirty minutes generating and stripping data. Put it away unread.
  • Week 2 — solve. A full paper as normal. Do not look at your draft.
  • Week 3 — sit your own. Solve your problem cold, timed. Then run the delete-a-row grid and the rival-analysis check, and write down every rule you could not state cleanly.
  • Week 4 — trade. Swap problems with someone in your club and mark each other’s. Watching a real solver misread your data is the fastest feedback in the whole method.

The list of rules you could not state cleanly, accumulated over a term, is the most honest revision list you will ever have — better than any topic checklist, because you generated it from your own failures rather than from a syllabus. Slot the rotation alongside our twelve-week study plan rather than in place of it.

Questions we get asked

Should I use a real language or invent one?
Invent one at first. Real languages carry irregularities you did not choose, which makes the sufficiency test much harder to run cleanly.

How many rules should a first attempt have?
Two that interact. One is trivial; three or more usually produces accidental ambiguity you will not spot.

Is setting problems a substitute for solving them?
No. Solving is the skill being tested. We suggest roughly four hours of solving to one of setting.

Can I reuse data from a published competition paper?
Study the structure, not the content. Copying data teaches nothing and is not yours to republish.

This is an independent guide operated by Hanlin Education for China-based international-school students. We are not affiliated with, endorsed by, or sponsored by the IOL Board. The language system used above is invented for this article and reproduces no competition material; contest formats and regulations are set by the organisers and can change — confirm current details on ioling.org. Errors reported to our editorial desk are corrected within 7 working days.

Scan · IOL Club mentor

Not sure where to start?

One WhatsApp message to the club mentor — ask about:

  • Registration help for your national olympiad
  • Free past-paper archive (22 years) + solutions
  • Weekly walkthroughs + a mentor for your grade
Scan the WeChat QR code
WhatsApp
WhatsApp
WeChat
WeChat

Name + school year + country

More from the club

Keep reading.

All bulletins →