Marzena Karpinska
Assistant Professor at Simon Fraser University
prospective students
I plan to take ~1 PhD student every year in the Fall. Please apply through the standard SFU process. There is no need to email me. If you would like to work with me on a project, fill in this form. SFU students can also join through CMPT 415/416. Please note that any AI-generated email will be deleted without consideration.
More details
Can you consult on my application or answer my email?
Sadly, no. Due to the volume of emails I receive, I cannot consult on applications or respond to individual emails from prospective students. I check the form periodically and will get back to you if there is a good fit.
Master's or PhD?
SFU, like other Canadian universities, offers separate Master's and PhD programs. I may need to prioritize PhD applications and accept Master's students only when funding permits. You can also apply directly to a PhD program.
What is CMPT 415/416?
It is a research course that lets you do a research project for credit over two semesters. You will meet with me weekly and report your progress. We work as a team, and while it is a lot of work, it is also a lot of fun. You can also help with research without taking the course if you are motivated and serious about it.
about me
I am currently an assistant professor at the Simon Fraser University in beautiful Vancouver, Canada. Before that, I was a senior researcher at Microsoft based in Redmond. I did my postdoc at the Manning College of Information & Computer Sciences, University of Massachusetts Amherst working with Prof. Mohit Iyyer.
I hold a Ph.D. from the Department of Language and Information Sciences at the University of Tokyo.
research
My research focuses on natural language processing (NLP) and language models (LMs). Specifically, I am interested in:
- Detection of AI-generated and AI-translated text: what characterizes text that is generated or translated by AI, and how well can we distinguish it from human-authored text?
- Multilingual NLP: how can we build systems that serve people from diverse linguistic and cultural backgrounds?
- Machine translation of creative texts: how can we help more authors be heard without misrepresenting their voice?
- Long-form text processing: how well can models process long-form input, and how can we evaluate long-form output?
media
mentions refers to our findings · discusses features our work · quotes quoted
- EL PAÍS
- RADIO-CANADA
- NATURE
- FUTURISM
- THE ATLANTIC
- WSJ
- PRESS GAZETTE
- FORBES
- THE ECONOMIST
- TECH CRUNCH
- UMASS MAGAZINE
- SLATOR
news
- Aug 2026 Our paper on long-context literary machine translation was accepted to EMNLP Findings 🎉
- Aug 2026 Our lab received a 5-year NSERC Discovery Grant & Launch Supplement 🎉
- Apr 2026 Our paper on AI-generated text in US news articles was accepted to ACL 2026 (oral, 3.9% papers) 🎉
- Jan 2026 Started as an assistant professor at Simon Fraser University in Vancouver, Canada.
- Oct 2025 New paper on AI-generated text in US news articles is out
- Aug 2025 Papers on interdisciplinary approach to MT, cross-lingual memorization, and quantization effects on long-context tasks accepted to EMNLP 2025 🎉
- Jul 2025 Our work on multilingual long-context processing was accepted to COLM 🎉
- Jun 2025 Preprint on interdisciplinary approach to machine translation is out
- Jun 2025 Serving as a Senior Area Chair at EMNLP
- May 2025 Our preprint on cross-lingual memorization is out
- May 2025 Our preprint on the effect of quantization on long-context tasks is out
- May 2025 Our works on AI-generated text and multilingual long-form QA were accepted to ACL 🎉
- Mar 2025 Our preprint on multilingual long-context processing is out
- Jan 2025 Our preprint on machine generated text detection is out
- Jan 2025 New preprint on slow-down attacks on reasoning models is out
- Nov 2024 Recognized as an outstanding AC at EMNLP 2024
- Sep 2024 NoCha was accepted to EMNLP 2024 🎉
- Sep 2024 ESA was accepted to WMT 2024 🎉
- Aug 2024 Presented our work on the evaluation of long-context language models at UNSW
- Jul 2024 Presented our work on the evaluation of long-context language models at RMIT
- Jul 2024 Presented our work on the evaluation of long-context language models at the University of Melbourne
- Jul 2024 FABLES was accepted to COLM 2024 🎉
- Jun 2024 Our preprint on LONG-CONTEXT processing capabilities of language models is out
- Jun 2024 Our preprint on MULTI-LINGUAL/CULTURAL performance of language models is out
- Jun 2024 Our preprint on more robust evaluation for machine translation is out
- Apr 2024 Our preprint on faithfulness in book-length summaries is out
- Mar 2024 NarrativeTime was accepted to LREC-COLING 2024 🎉
- Dec 2023 Presented our work at WMT in Singapore.
- Nov 2023 Launched litmt.org, a platform for sharing machine-translated world literature.
- May 2023 Virtual talk at Instituto Superior Técnico & Unbabel Seminar on translation with Large Language Models.
- Apr 2023 Virtual talk at Microsoft MT Reading Group on translation with Large Language Models.
- Jan 2023 Virtual talk at Polish Academy of Sciences on Evaluation of Long-form Text Generation.
- Dec 2022 Presented our work on diagnosing automatic evaluation metrics at EMNLP in Abu Dhabi.
- Dec 2022 Presented our work on document-level MT at EMNLP in Abu Dhabi.
- Nov 2022 Slator on literary MT and our PAR3 work