Morphology in LLMs
Bachelor’s Thesis on Language Models
2026 · Bachelor’s thesis, TUM
A Comparative Study of Morphological Capabilities in Language Models Across Agglutinative Languages
Take apart
Segmentation
Split a Korean word into its pieces and name each one.
Edit
Case transformation
Change one grammatical ending and keep the rest.
Build
Generation
Form a word from its dictionary form and a list of features.
PROBLEM & USERS
Language models write fluent text, but does that mean they understand how words are built? In Finnish, Hungarian, Korean and Turkish, one word can stack several meanings: Turkish evlerimizden means “from our houses”. Most benchmarks score whole sentences, so a single wrong ending goes unnoticed. This thesis tests word grammar directly.
Who it’s for: Researchers and engineers building language tools for languages like these, who need to know what a model really gets right.
- evhouse
- lerplural
- imizour
- denfrom
evlerimizden“from our houses”
CONSTRAINTS
- Four languages: from three language families, written in two scripts
- Unseen words: the trained model is tested on words it never saw in training
- Strict scoring: an answer counts only if it matches the correct word exactly
- Limited compute: a small 3B model, trained in a memory-saving 4-bit mode (QLoRA)
APPROACH
- 01 Take apart도움을Korean, “help” as object도움을“help” + object marker
- 02 EditilişkilerdeTurkish, “in relations”ilişkilere“to relations”
- 03 Buildkatsella+ featuresFinnish, “to watch”katsellen“by watching”
Task 1: take a word apart
Split a Korean word into its pieces and name each piece. For example, 도움을 becomes 도움 (“help”) + 을 (object marker). 500 words.
Task 2: edit a word
Change only the grammatical case and keep everything else. For example, Turkish ilişkilerde (“in relations”) becomes ilişkilere (“to relations”). 2,000 words in four languages.
Task 3: build a word
Form a word from its dictionary form and a list of features. For example, Finnish katsella (“to watch”) becomes katsellen (“by watching”). 2,000 words in four languages.
Prompting, then training
First, ask existing models (GPT-4o, Gemini 3 Flash and Llama 3B, 8B and 70B), with and without five worked examples. Then train the smallest model on the task and compare it with its untrained self.
WHAT SHIPPED
A three-task test suite with datasets built from public language resources (Universal Dependencies and UniMorph), results for five models, and a trained version of Llama 3.2 3B. Defended at TUM on 28 September 2026, supervised by Dr. Marion Di Marco and examined by Prof. Dr. Alexander M. Fraser.
Architecture Snapshot
Input
Researchers and engineers building language tools for languages like these, who need to know what a model really gets right.
Core Decision
Test word grammar three ways: take apart, edit and build
Output
One score hides most of what a model can and cannot do
Stack
- Python
- PyTorch
- Transformers
- PEFT / QLoRA
- GPT-4o
- Gemini 3 Flash
- Llama 3
- Universal Dependencies
- UniMorph
IMPACT
Korean · five-shot · F1 score
- Pieces
- Pieces + labels
- GPT-4o
- Gemini3 Flash
Models find the pieces more reliably than they label them.
Five-shot · % correct
- Editing
- Building
- Llama3B
- Llama8B
- Llama70B
- GPT-4o
Building was measured on an earlier data split.
Average % correct on unseen words
- 3B base
- 3B + QLoRA
- GPT-4o
- Editing
- Building
GPT-4o reference: zero-shot on the same words for editing, five-shot on an earlier split for building.
- Taking apart: with five examples, GPT-4o finds the pieces of Korean words well (92.3 out of 100) but labels them less reliably (81.5)
- Editing: five worked examples help every model, by 11 to 36 points. GPT-4o is best at 94.3% correct
- Building: bigger models do better. With examples, the 3B model gets 16% right, 70B gets 63% and GPT-4o 79%
- Training: the small 3B model jumps from 12% to 84% correct on editing and from 7% to 79% on building, on words it never saw
LEARNINGS
- Word grammar is not one skill: a single score hides most of the story
- Following the task matters: many low scores came from models copying the input back, not from missing knowledge of the language
- Small can be enough: for a narrow task, a small trained model comes close to a much larger one and runs cheaply on your own hardware
- Tokens are not the whole story: how a model cuts text into pieces is linked to some errors, but it does not explain most of them
NEXT STEPS
- Made-up words, to see whether models apply rules or remember forms
- Training on three languages and testing on the fourth, to see how far learning transfers
- Models that read text letter by letter, to see whether token cutting is the bottleneck