Tamil Semantic Tokenizer
Morphology-Aware Semantic Tokenizer for Tamil based on FST models

Software engineering · AI research
AI-Native Builder, Engineer and Researcher
AI-Native Builder, Engineer and Researcher with more than a decade of experience building production software.
anand.murugan@gmail.comResearch and analysis

An experiment-by-experiment study of whether action-level feedback can improve reinforcement learning for language models.

A response to the LM-head bottleneck claim using backward-only rank controls, repeated-token exposure tests, and small language models.
A technical study of reversible Tamil semantic tokenization, hierarchical word composition, and a controlled six-system translation comparison.

How fewer than 140,000 root lemmas and inflectional rules generate more than 2.13 billion distinct Tamil written forms.

Experiments with entity registers and typed references for exact values—and why choosing the right entity remained hard.

An ontological exploration of duality and the fundamental relation between Word and Form.

பௌத்த தத்துவத்தின் மையக் கருத்து
Selected work
Morphology-Aware Semantic Tokenizer for Tamil based on FST models




Adaptive two-way practice for Tamil vowels, consonants, Grantha letters, and உயிர்மெய் forms.

Reading, transliteration, and meaning practice across 5,000 Tamil words.

Adaptive two-way practice matching Japanese hiragana with their romaji readings.

Reading, transliteration, and meaning practice across 5,000 Japanese words.


Multi-tenant commerce platform with merchant and customer applications.