Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretability

Abstract

This work pre-trains Tsetlin Machines with semantic clusters derived from language models, giving logic-based learners a semantically meaningful starting point. The resulting models remain interpretable by design while benefiting from the semantic knowledge captured by modern language models, improving accuracy on text classification without sacrificing transparency.

Publication
arXiv preprint
Yuangang Li
Yuangang Li
PhD Student at UCI