arXiv:2509.07274v4 Announce Type: replace-cross
Abstract: Migration has been a core topic in German political debate, from postwar expellee displacement to labor migration and recent refugee movements. Large-scale analysis of such political discourse has traditionally required extensive manual annotation, limiting coverage. Large language models (LLMs) offer a scalable alternative. Using a theory-driven annotation scheme, we examine how well LLMs annotate subtypes of solidarity and anti-solidarity in German parliamentary debates and whether the resulting labels support valid downstream inference. We first evaluate multiple LLMs across model size, prompting strategies, fine-tuning, historical versus contemporary data, and systematic errors. The strongest models, especially GPT-5 and gpt-oss-120B, achieve macro- F1 scores comparable to human agreement, although their systematic errors can bias downstream results. We therefore combine soft-label model outputs with Design-based Supervised Learning (DSL) to reduce bias in long-term trend estimates. Beyond the methodological evaluation, we interpret the resulting annotations from a social-scientific perspective across the Reichstag (1867- 1933), West German Bundestag (1949-1990) and Re-Unified German Bundestag (1990-2025) corpus to trace trends in solidarity and anti-solidarity toward migrants, with detailed analysis of postwar Germany (1949-1957) and contemporary Germany (2009-2025). We find relatively high levels of solidarity in the postwar period, especially in group-based and compassionate forms, and a marked rise in anti-solidarity since 2015. We argue that LLMs can support large-scale social-scientific text analysis, but their outputs require rigorous validation and, where systematic errors are present, statistical correction.
