Second Tokenization Workshop
Tomasz Limisiewicz ⋅ Valentin Hofmann ⋅ Sachin Kumar ⋅ Jindřich Libovický ⋅ Jindřich Helcl ⋅ Orevaoghene Ahia ⋅ Elizabeth Salesky ⋅ Yuval Pinter ⋅ Yuki M Asano
Abstract
Tokenization–the process of converting raw data into discrete units for model input and output–has emerged as a critical component across machine learning domains. Originally central to natural language processing (NLP), tokenization is now equally essential in multimodal learning, computer vision, speech processing, and other areas. Recent research has shown that tokenization strategies significantly impact model utility, efficiency, and generalization, sparking a surge of interest in this foundational topic.
Schedule
Time zone: Conference local time zone
Successful Page Load