About me

Hi, I am Tomasz Limisiewicz. I am currently exploring large language models as a Postdoctoral Researcher at the University of Washington and Meta in Seattle. I PhDid at Charles University.

How do language models perform so well across diverse tasks, and how do their capabilities transfer across languages and modalities? These questions are core to my research. In my exploration, I focus on the following principles:

I. Focus on Data and Tokens

Studying corpus artifacts and phenomena characteristic of specific languages and modalities is crucial to understanding how language models work. I am especially interested in the role of tokenization, which mediates between data and model.

II. Divide and Explain

I dissect black-box models and analyze specific components: attention, latent representations, and tokenization. This approach helps identify the role of architectural choices in training and explain which changes bring improvements at scale.

III. Transfer and Robustness

Joint training across domains enables knowledge transfer and faster learning, for example in languages with limited data. My aim is to build systems that work across languages and modalities and improve together by learning similar, transferable representations.

For a selection of my work, see my publications; for a full list, see my Google Scholar profile.

Beyond research, I’m a keen hiker and a film enthusiast.