Qualitative Research on Quantitative Scale? Opportunities and Pitfalls of LLMs as Annotators
Many TRR researchers had the experience of having to painstakingly annotate a long transcript by hand. Due to the time needed to perform such deep and detail-oriented work, such kind of analysis is not easily applicable to larger data sets. Large language models now seem to enable us to do just that: prompting an LLM with an complex annotation guideline and then simply applying the model to an arbitrarily large data set, thus seemingly scaling up qualitative research. Using the example of a multi-year project on annotating (anti-)solidarity towards migrants in German parliamentary speeches and the example of a large-scale semi-automated literature review, we'll discuss the opportunities and the (many) menthodological pitfalls of using LLMs as annotators.
References:
Kostikova, Aida, Wang, Zhipin, Bajri, Deidamea, Pütz, Ole, Paaßen, Benjamin, Eger, Steffen (2026). LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models. ACM Computing Surveys. doi:10.1145/3801096
Aida Kostikova, Ole Pütz, Steffen Eger, Olga Sabelfeld, Benjamin Paassen (2025). LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade. arXiv. https://arxiv.org/abs/2509.07274
Speaker information:
Benjamin Paaßen is junior professor for Knowledge Representation and Machine Learning at Bielefeld University and Co-PI in project C01 on Healthy Distrust. Their research foci are machine learning for education, explainable and interpretable machine learning, as well as limitations of large language models.
Online, via Zoom