Project Description
Recent LLMs have demonstrated remarkable capabilities across a range of tasks, yet their underlying linguistic competence remains less explored. As these models increasingly inform research and applications in the social sciences and beyond, a systematic evaluation of their linguistic abilities becomes critical. This project investigates the extent to which LLMs capture core aspects of linguistic knowledge, including traditional categories in syntax, semantics, and pragmatics, as well as phenomena studied in sociolinguistic variation and psycho-, neuro-, and biolinguistics. Drawing on methods from linguistics, cognitive science, and natural language processing, we design targeted evaluation tasks to probe these specific linguistic phenomena, develop new benchmarks, and identify systematic strengths and weaknesses in current models. Our aim is to contribute to a more rigorous understanding of LLM capabilities and limitations, providing insights that are essential for both theoretical modeling and practical deployment.
For master/bachelor students: if you are interested in writing a thesis on this topic, please feel free to reach out to me (bolei.ma@lmu.de) directly.
Contact person
Related Publications
- Bolei Ma. 2024. Evaluating Lexical Aspect with Large Language Models. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 123–131, Bangkok, Thailand. Association for Computational Linguistics.
- Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, and Barbara Plank. 2025. Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8679–8696, Vienna, Austria. Association for Computational Linguistics.
- Yu Lei, Xingyang Ge, Yi Zhang, Yiming Yang, and Bolei Ma. 2026. Do Large Language Models Think like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRI. Proceedings of the AAAI Conference on Artificial Intelligence, 40(1), 579–587.
- Bolei Ma and Yusuke Miyao. 2026. The Imperfective Paradox in Large Language Models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15093–15111, San Diego, California, United States. Association for Computational Linguistics. 🏆 Best Paper Award @ ACL, see here.
- Yajie Wen, Ziwei Gong, Chengyan Wu, Xiyun Gong, Yun Xue, Julia Hirschberg, and Bolei Ma. 2026. YUE-PUB-Speech: A Speech-based Pragmatic Understanding Benchmark for Cantonese. Interspeech.