Morphology in LLMs

Bachelor’s Thesis on Language Models

2026 · Bachelor’s thesis, TUM

LLMNLPQLoRAPython

A Comparative Study of Morphological Capabilities in Language Models Across Agglutinative Languages

Take apart

Segmentation

Split a Korean word into its pieces and name each one.

Edit

Case transformation

Change one grammatical ending and keep the rest.

Build

Generation

Form a word from its dictionary form and a list of features.

Language models write fluent text, but does that mean they understand how words are built? In Finnish, Hungarian, Korean and Turkish, one word can stack several meanings: Turkish evlerimizden means “from our houses”. Most benchmarks score whole sentences, so a single wrong ending goes unnoticed. This thesis tests word grammar directly.

Who it’s for: Researchers and engineers building language tools for languages like these, who need to know what a model really gets right.

One Turkish word, four meanings
  1. evhouse
  2. lerplural
  3. imizour
  4. denfrom

evlerimizden“from our houses”

  • Four languages: from three language families, written in two scripts
  • Unseen words: the trained model is tested on words it never saw in training
  • Strict scoring: an answer counts only if it matches the correct word exactly
  • Limited compute: a small 3B model, trained in a memory-saving 4-bit mode (QLoRA)
Three ways to test a word
  1. 01 Take apart도움을Korean, “help” as object도움을“help” + object marker
  2. 02 EditilişkilerdeTurkish, “in relations”ilişkilere“to relations”
  3. 03 Buildkatsella+ featuresFinnish, “to watch”katsellen“by watching”

Task 1: take a word apart

Split a Korean word into its pieces and name each piece. For example, 도움을 becomes 도움 (“help”) + 을 (object marker). 500 words.

Task 2: edit a word

Change only the grammatical case and keep everything else. For example, Turkish ilişkilerde (“in relations”) becomes ilişkilere (“to relations”). 2,000 words in four languages.

Task 3: build a word

Form a word from its dictionary form and a list of features. For example, Finnish katsella (“to watch”) becomes katsellen (“by watching”). 2,000 words in four languages.

Prompting, then training

First, ask existing models (GPT-4o, Gemini 3 Flash and Llama 3B, 8B and 70B), with and without five worked examples. Then train the smallest model on the task and compare it with its untrained self.

A three-task test suite with datasets built from public language resources (Universal Dependencies and UniMorph), results for five models, and a trained version of Llama 3.2 3B. Defended at TUM on 28 September 2026, supervised by Dr. Marion Di Marco and examined by Prof. Dr. Alexander M. Fraser.

Architecture Snapshot

Input

Researchers and engineers building language tools for languages like these, who need to know what a model really gets right.

Core Decision

Test word grammar three ways: take apart, edit and build

Output

One score hides most of what a model can and cannot do

Stack

  • Python
  • PyTorch
  • Transformers
  • PEFT / QLoRA
  • GPT-4o
  • Gemini 3 Flash
  • Llama 3
  • Universal Dependencies
  • UniMorph
Taking apart

Korean · five-shot · F1 score

  • Pieces
  • Pieces + labels
  • Pieces: 92.3Pieces + labels: 81.5GPT-4o
  • Pieces: 94.9Pieces + labels: 88.5Gemini3 Flash

Models find the pieces more reliably than they label them.

Prompting

Five-shot · % correct

  • Editing
  • Building
  • Editing: 47.8Building: 16.2Llama3B
  • Editing: 70Building: 37.4Llama8B
  • Editing: 78.3Building: 62.8Llama70B
  • Editing: 94.3Building: 79GPT-4o

Building was measured on an earlier data split.

Training the 3B model

Average % correct on unseen words

  • 3B base
  • 3B + QLoRA
  • GPT-4o
  • 3B base: 12.43B + QLoRA: 83.6GPT-4o: 83.2Editing
  • 3B base: 7.13B + QLoRA: 78.6GPT-4o: 79Building

GPT-4o reference: zero-shot on the same words for editing, five-shot on an earlier split for building.

  • Taking apart: with five examples, GPT-4o finds the pieces of Korean words well (92.3 out of 100) but labels them less reliably (81.5)
  • Editing: five worked examples help every model, by 11 to 36 points. GPT-4o is best at 94.3% correct
  • Building: bigger models do better. With examples, the 3B model gets 16% right, 70B gets 63% and GPT-4o 79%
  • Training: the small 3B model jumps from 12% to 84% correct on editing and from 7% to 79% on building, on words it never saw
  • Word grammar is not one skill: a single score hides most of the story
  • Following the task matters: many low scores came from models copying the input back, not from missing knowledge of the language
  • Small can be enough: for a narrow task, a small trained model comes close to a much larger one and runs cheaply on your own hardware
  • Tokens are not the whole story: how a model cuts text into pieces is linked to some errors, but it does not explain most of them
  • Made-up words, to see whether models apply rules or remember forms
  • Training on three languages and testing on the fourth, to see how far learning transfers
  • Models that read text letter by letter, to see whether token cutting is the bottleneck
Next projectWortschaftGerman Vocabulary Learning App