Corpus Data

The full corpus, consisting of full-length audio, textgrids, transcriptions, metadata, and a pre-generated phoneme-level acoustic dataset (via the Praat-Mass-Analyzer ) for all 603 speakers, will release once the research team completes its testing phase.

In the meantime, a sizable sample of the corpus, consisting of the aforementioned files for 90 speakers, is available upon email request. For access to this MuHSiC sample, contact muhsic at berkeley.edu

For users that previously requested and received access, the sample is hosted here .

The full corpus contains 603 speakers and 716 hours of audio, with 35 minutes of English sociolinguistic interview audio and 35 minutes of Spanish sociolinguistic interview audio per speaker. Metadata has been digitized into a csv format using the interview materials below. All audio files are available in 44.1 kHZ, 16-bit wav or mp3 format. All audio has been transcribed in pdf and docx format, and force aligned using the Montreal Forced Aligner (TextGrid format).

If you do not wish to download the corpus, you may search truncated, 5-minute recordings from the public sample of 90 speakers below.

Interview Materials

MuHSiC Interviewer Manual
English Sociolinguistic Interview Questions
Spanish Sociolinguistic Interview Questions
Interviewer Background Form
Interviewer Field Notes
Cuestionario Sociodemográfico
BLP (English–Spanish)
BLP (Spanish–English)

Search and listen to a sample of corpus audio

Search every column at once or select one column to search within. Both searches can be used together. Click on a row to listen to the audio and read the transcript.

Loading corpus data…

0 matching records. This is a sample of available metadata. For full metadata access, please request access.

Loading corpus data…