Adaptive Keyword Extraction Service for Turkish
Tarih
Yazarlar
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Erişim Hakkı
Özet
Keyword extraction is one of the important tasks of NLP. In this study, we have implemented a fast and competitive keyword extraction model for Turkish. We have benchmarked our model with all available online keyword extraction models such as RAKE, YAKE, and TF/IDF. We have used fastText for word embeddings which produced well-distributed clusters. For obtaining keywords, we have developed a dynamic cluster selection mechanism that can vary through input text. The keyword selection mechanism is based on word votes on input text which detects target cluster(s) to keyword set. Every word has a unique vote according to its IDF score. We also established a penalty function for weakening verbs. Our model enforces the constraint that common verbs in Turkish should not be selected as a keyword candidate. Our model can infer keywords in a millisecond timeframe. It does not require GPU or TPU for inference. It is a cost-effective, online, self-learning, and fast keyword extraction model.








