Swiss-AL Corpora#
The Swiss-AL corpus collection supports discourse-analytical research across Swiss public communication. It comprises organizational, journalistic, and parliamentary corpora, together with project-specific corpora developed by the ZHAW Digital Discourse Lab.
Important
You need to log in and create a project to access the full collection. Without an account you can only use the demo corpora.
Corpus types#
The collection is organized into types. A corpus’s type is not just a label: it determines how the corpus sources are subcategorized, and therefore which categories are available when you group results in Distribution of Words.
Type |
What it contains |
Sources subcategorized by |
|---|---|---|
Web data from actors in Swiss public communication |
social system, field of activity |
|
Swiss news media |
publisher, media type |
|
Swiss Federal Parliament debates |
— |
|
Collections built for specific research projects |
varies by corpus |
|
Subsets of the organizational corpora, open without login |
social system, field of activity |
Note
Right-wing extremist corpora are planned but not yet available. This category will include texts from far-right news platforms such as PI-News and Compact, for research into the spread of political ideologies, rhetoric, and narratives of the far right. For a case study using this data, see Krasselt et al. 2022.
All corpora at a glance#
Corpus |
Language |
Period |
Documents |
Tokens |
Type |
DOI |
|---|---|---|---|---|---|---|
German |
2010–2024 |
607,040 |
265,753,260 |
Organizational |
— |
|
French |
2010–2024 |
212,492 |
117,372,532 |
Organizational |
— |
|
Italian |
2010–2024 |
147,788 |
71,812,596 |
Organizational |
— |
|
German |
2018–2025 |
1,426,368 |
787,564,272 |
Journalistic |
||
German |
2010–2025 |
2,384,267 |
1,206,855,045 |
Journalistic |
— |
|
French |
2018–2025 |
1,911,560 |
1,158,001,810 |
Journalistic |
— |
|
French |
2010–2025 |
2,358,649 |
1,308,104,934 |
Journalistic |
— |
|
Italian |
2018–2025 |
625,532 |
255,109,260 |
Journalistic |
— |
|
Italian |
2010–2025 |
154,998 |
64,542,198 |
Journalistic |
— |
|
Rumantsch |
2018–2025 |
108,222 |
44,177,318 |
Journalistic |
— |
|
German |
2000–2024 |
806,772 |
61,067,496 |
Parliamentary |
— |
|
French |
2000–2024 |
284,232 |
23,605,674 |
Parliamentary |
— |
|
Italian |
2000–2024 |
11,034 |
899,682 |
Parliamentary |
— |
|
German |
2014–2024 |
6,262 |
6,665,202 |
Project |
— |
|
German |
2000–2025 |
603,554 |
490,248,554 |
Project |
— |
|
German |
2000–2024 |
27,378 |
19,074,618 |
Project |
— |
|
French |
2000–2025 |
1,911,560 |
1,158,001,810 |
Project |
— |
|
Italian |
2000–2025 |
1,764 |
1,265,758 |
Project |
— |
|
German |
2000–2025 |
780,247 |
629,479,591 |
Project |
— |
|
German |
2000–2024 |
9,048 |
5,077,906 |
Project |
— |
|
German |
— |
22,983 |
10,029,710 |
Demo |
— |
|
French |
— |
18,636 |
10,061,634 |
Demo |
— |
|
Italian |
— |
20,500 |
10,050,082 |
Demo |
— |
Update frequency#
Organizational and journalistic corpora are updated annually. Parliamentary, project, and demo corpora are static unless otherwise noted.
Creating subcorpora#
You can build your own subcorpus by selecting Create custom corpus and using the wizard to choose categories, time span, and sources. You can also build a subcorpus around a specific lemma by entering it as a keyword — for example, entering Corona collects documents containing that word.
Keyword filtering is case-insensitive and matches all forms in the word’s paradigm, exactly as in Quick Search.
Tip
Use only single-word keywords; multiword expressions are not supported.
With several keywords, ANY returns documents containing at least one of them (OR), while ALL returns only documents containing every one of them (AND).
Important
Subcorpora must meet minimum sizes for some tools:
Topic modelling requires at least 1,000 documents.
Word embeddings, used by Semantic Space, require at least 10 million tokens.
Check these in Overview after building a subcorpus.