Ensemble models
MalayBERT macro F1
Pantun theme classes
Malay pantun has a two-part structure: the pembayang sets an image, the maksud carries the point. Feed a classifier both halves and the imagery drowns the meaning. This work segments them first, then runs three very different models over what remains (a fine-tuned MalayBERT, a TF-IDF + SVM, and a TextCNN), resolving their predictions through a majority vote.
The problem
A pantun’s first couplet is deliberately oblique: it exists to set rhyme and rhythm, not to state the theme. A bag-of-words model trained on the full quatrain therefore learns from text that is, by design, not about the subject. Malay also lacks the mature preprocessing tooling English enjoys, and the corpus is small enough that vocabulary sparsity becomes the dominant failure mode rather than an afterthought.
What I built
A regex-based segmenter isolates maksud from pembayang before anything else runs, followed by case folding, tokenisation and Porter stemming to pull the vocabulary down to a size the corpus can actually support. Three models are then trained on identical splits and played to their strengths: a fine-tuned MalayBERT reads whole-context meaning and figurative intent; an RBF-kernel SVM over 10,000-dimensional TF-IDF unigram and bigram features catches explicit keyword signals; a PyTorch TextCNN captures local n-gram patterns. A majority-vote layer resolves the three into a single consensus theme with per-model confidence.
How it works
Pembayang/maksud segmentation
A regex segmenter splits the quatrain on its structural boundary so the classifier trains on the semantically loaded half. This is the step that makes the rest of the pipeline worth running.
Normalisation for a small corpus
Case folding, tokenisation and Porter stemming, chosen specifically to reduce vocabulary sparsity. On a corpus this size, unchecked vocabulary growth hurts more than the information lost to stemming.
SVM baseline
RBF-kernel SVM over 10,000-dimensional TF-IDF unigram and bigram features with sublinear term-frequency scaling: a strong, honest baseline that a neural model has to actually beat.
TextCNN
Parallel 1D convolutional layers at several kernel widths with global max-pooling, capturing localised n-gram syntax that neither the transformer nor the bag-of-words baseline picks up the same way.
MalayBERT fine-tune
mesolitica/bert-base-standard-bahasa-cased fine-tuned on the segmented maksud. It reads whole-context meaning and figurative intent, which is precisely where keyword-driven approaches collapse on pantun.
Majority-vote ensemble
Rather than picking one winner, the three predictions resolve through a majority vote that reports per-model confidence alongside the consensus theme: an explainable read rather than a black-box label.
Deployment
A Flask API that lazy-loads weights (joblib, pth, safetensors) on first request and caches them in process, holding response times under one second without keeping every model resident at boot.
Outcome
MalayBERT led on nuance at ~60% macro F1, with SVM close behind (~55%) and TextCNN trailing (~47%), exactly as the data scarcity predicted. The transparent three-model view turns a black-box label into an explainable read of each verse, deployed behind a Flask API that lazy-loads and caches weights to keep inference under one second. Accepted to AiDAS 2026.
Stack
