Skip to main navigation Skip to search Skip to main content

GRASS: The Graz Corpus of Read and Spontaneous Speech

Research output: Chapter in Book/Report/Conference proceedingConference paperpeer-review

Abstract

This paper provides a description of the preparation, the speakers, the recordings, and the creation of the orthographic transcriptions of the first large scale speech database for Austrian German. It contains approximately 1900 minutes of (read and spontaneous) speech
produced by 38 speakers. The corpus consists of three components. First, the Conversation Speech (CS) component contains free conversations of one hour length between friends, colleagues, couples, or family members. Second, the Commands Component (CC)
contains commands and keywords which were either read or elicited by pictures. Third, the Read Speech (RS) component contains phonetically balanced sentences and digits. The speech of all components has been recorded at super-wideband quality in a soundproof
recording-studio with head-mounted microphones, large-diaphragm microphones, a laryngograph, and with a video camera. The orthographic transcriptions, which have been created and subsequently corrected manually, contain approximately 290 000 word tokens
from 15 000 different word types.
Original languageEnglish
Title of host publication9th Conference on Language Resources and Evaluation Conference (LREC 2014)
Place of PublicationRed Hook, NY
PublisherCurran
Pages1465-1470
Volume2
ISBN (Print) 9781632666215
Publication statusPublished - 2014
Event9th International Conference on Language Resources and Evaluation: LREC 2014 - Reykjavik, Iceland
Duration: 26 May 201431 May 2014

Conference

Conference9th International Conference on Language Resources and Evaluation
Abbreviated titleLREC 2014
Country/TerritoryIceland
CityReykjavik
Period26/05/1431/05/14

Fields of Expertise

  • Information, Communication & Computing

Treatment code (Nähere Zuordnung)

  • Application
  • Basic - Fundamental (Grundlagenforschung)

Fingerprint

Dive into the research topics of 'GRASS: The Graz Corpus of Read and Spontaneous Speech'. Together they form a unique fingerprint.

Cite this