Projects per year
Abstract
This paper provides a description of the preparation, the speakers, the recordings, and the creation of the orthographic transcriptions of the first large scale speech database for Austrian German. It contains approximately 1900 minutes of (read and spontaneous) speech
produced by 38 speakers. The corpus consists of three components. First, the Conversation Speech (CS) component contains free conversations of one hour length between friends, colleagues, couples, or family members. Second, the Commands Component (CC)
contains commands and keywords which were either read or elicited by pictures. Third, the Read Speech (RS) component contains phonetically balanced sentences and digits. The speech of all components has been recorded at super-wideband quality in a soundproof
recording-studio with head-mounted microphones, large-diaphragm microphones, a laryngograph, and with a video camera. The orthographic transcriptions, which have been created and subsequently corrected manually, contain approximately 290 000 word tokens
from 15 000 different word types.
produced by 38 speakers. The corpus consists of three components. First, the Conversation Speech (CS) component contains free conversations of one hour length between friends, colleagues, couples, or family members. Second, the Commands Component (CC)
contains commands and keywords which were either read or elicited by pictures. Third, the Read Speech (RS) component contains phonetically balanced sentences and digits. The speech of all components has been recorded at super-wideband quality in a soundproof
recording-studio with head-mounted microphones, large-diaphragm microphones, a laryngograph, and with a video camera. The orthographic transcriptions, which have been created and subsequently corrected manually, contain approximately 290 000 word tokens
from 15 000 different word types.
| Original language | English |
|---|---|
| Title of host publication | 9th Conference on Language Resources and Evaluation Conference (LREC 2014) |
| Place of Publication | Red Hook, NY |
| Publisher | Curran |
| Pages | 1465-1470 |
| Volume | 2 |
| ISBN (Print) | 9781632666215 |
| Publication status | Published - 2014 |
| Event | 9th International Conference on Language Resources and Evaluation: LREC 2014 - Reykjavik, Iceland Duration: 26 May 2014 → 31 May 2014 |
Conference
| Conference | 9th International Conference on Language Resources and Evaluation |
|---|---|
| Abbreviated title | LREC 2014 |
| Country/Territory | Iceland |
| City | Reykjavik |
| Period | 26/05/14 → 31/05/14 |
Fields of Expertise
- Information, Communication & Computing
Treatment code (Nähere Zuordnung)
- Application
- Basic - Fundamental (Grundlagenforschung)
Fingerprint
Dive into the research topics of 'GRASS: The Graz Corpus of Read and Spontaneous Speech'. Together they form a unique fingerprint.Projects
- 1 Finished
-
CLCS - Cross-layer pronunciation modeling for conversational speech
Schuppler, B. (Project manager)
1/09/12 → 30/04/17
Project: Research project
-
Turn-taking annotation for quantitative and qualitative analyses of conversation
Kelterer, A. & Schuppler, B., 2025, In: arXiv.org e-Print archive. cs.CL, p. 1 41 p.Research output: Contribution to journal › Article
Open Access -
10 Years of GRASS development: Experiences from annotating a large corpus of conversational Austrian German
Schuppler, B., Kelterer, A. & Hagmüller, M., 2023.Research output: Contribution to conference › Abstract
Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS