Swiss-AL Corpora#

The Swiss-AL corpus collection supports discourse-analytical research across Swiss public communication. It comprises organizational, journalistic, and parliamentary corpora, together with project-specific corpora developed by the ZHAW Digital Discourse Lab.

Important

You need to log in and create a project to access the full collection. Without an account you can only use the demo corpora.

Corpus types#

The collection is organized into types. A corpus’s type is not just a label: it determines how the corpus sources are subcategorized, and therefore which categories are available when you group results in Distribution of Words.

Type

What it contains

Sources subcategorized by

Organizational

Web data from actors in Swiss public communication

social system, field of activity

Journalistic

Swiss news media

publisher, media type

Parliamentary

Swiss Federal Parliament debates

Project

Collections built for specific research projects

varies by corpus

Demo

Subsets of the organizational corpora, open without login

social system, field of activity

Note

Right-wing extremist corpora are planned but not yet available. This category will include texts from far-right news platforms such as PI-News and Compact, for research into the spread of political ideologies, rhetoric, and narratives of the far right. For a case study using this data, see Krasselt et al. 2022.

All corpora at a glance#

Corpus

Language

Period

Documents

Tokens

Type

DOI

DE Organizational Corpus

German

2010–2024

607,040

265,753,260

Organizational

FR Organizational Corpus

French

2010–2024

212,492

117,372,532

Organizational

IT Organizational Corpus

Italian

2010–2024

147,788

71,812,596

Organizational

DE Journalistic Corpus (high reach+regional, 2018-2025)

German

2018–2025

1,426,368

787,564,272

Journalistic

https://doi.org/10.48656/v776-mk85

DE Journalistic Corpus (high reach media, 2010-2025)

German

2010–2025

2,384,267

1,206,855,045

Journalistic

FR Journalistic Corpus (high reach+regional, 2018-2025)

French

2018–2025

1,911,560

1,158,001,810

Journalistic

FR Journalistic Corpus (high reach media, 2010-2025)

French

2010–2025

2,358,649

1,308,104,934

Journalistic

IT Journalistic Corpus (high reach+regional, 2018-2025)

Italian

2018–2025

625,532

255,109,260

Journalistic

IT Journalistic Corpus (high reach, 2010-2025)

Italian

2010–2025

154,998

64,542,198

Journalistic

RM Journalistic Corpus (high reach+regional, 2018-2025)

Rumantsch

2018–2025

108,222

44,177,318

Journalistic

DE Swiss Federal Parliament Debates Corpus (2000-2024)

German

2000–2024

806,772

61,067,496

Parliamentary

FR Swiss Federal Parliament Debates Corpus (2000-2024)

French

2000–2024

284,232

23,605,674

Parliamentary

IT Swiss Federal Parliament Debates Corpus (2000-2024)

Italian

2000–2024

11,034

899,682

Parliamentary

DE Journalistic Delinquency Corpus (2014-2024)

German

2014–2024

6,262

6,665,202

Project

DE Journalistic Sepsis Corpus (2000-25)

German

2000–2025

603,554

490,248,554

Project

DE Organizational Sepsis Corpus (2000-24)

German

2000–2024

27,378

19,074,618

Project

FR Journalistic Sepsis Corpus (2000-25)

French

2000–2025

1,911,560

1,158,001,810

Project

IT Journalistic Sepsis Corpus (2000-25)

Italian

2000–2025

1,764

1,265,758

Project

DE Journalistic VaxMo Corpus

German

2000–2025

780,247

629,479,591

Project

DE Journalistic VaxMo Corpus (keyword based, 2000-25)

German

2000–2024

9,048

5,077,906

Project

DE Demo Corpus

German

22,983

10,029,710

Demo

FR Demo Corpus

French

18,636

10,061,634

Demo

IT Demo Corpus

Italian

20,500

10,050,082

Demo

Update frequency#

Organizational and journalistic corpora are updated annually. Parliamentary, project, and demo corpora are static unless otherwise noted.

Creating subcorpora#

You can build your own subcorpus by selecting Create custom corpus and using the wizard to choose categories, time span, and sources. You can also build a subcorpus around a specific lemma by entering it as a keyword — for example, entering Corona collects documents containing that word.

Keyword filtering is case-insensitive and matches all forms in the word’s paradigm, exactly as in Quick Search.

Tip

Use only single-word keywords; multiword expressions are not supported.

With several keywords, ANY returns documents containing at least one of them (OR), while ALL returns only documents containing every one of them (AND).

Important

Subcorpora must meet minimum sizes for some tools:

  • Topic modelling requires at least 1,000 documents.

  • Word embeddings, used by Semantic Space, require at least 10 million tokens.

Check these in Overview after building a subcorpus.