The full corpus, consisting of full-length audio, textgrids, transcriptions, metadata, and a pre-generated phoneme-level acoustic dataset (via the
Praat-Mass-Analyzer
) for all 603 speakers, will release once the research team completes its testing phase.
In the meantime, a sizable sample of the corpus, consisting of the aforementioned files for 90 speakers, is available upon email request. For access to this MuHSiC sample, contact
muhsicberkeley.edu
For users that previously requested and received access, the sample is hosted
here
.
The full corpus contains 603 speakers and 716 hours of audio, with 35 minutes of English sociolinguistic interview audio and 35 minutes of Spanish sociolinguistic interview audio per speaker. Metadata has been digitized into a csv format using the interview materials below. All audio files are available in 44.1 kHZ, 16-bit wav or mp3 format. All audio has been transcribed in pdf and docx format, and force aligned using the Montreal Forced Aligner (TextGrid format).
If you do not wish to download the corpus, you may search truncated, 5-minute recordings from the public sample of 90 speakers below.
Search every column at once or select one column to search within.
Both searches can be used together. Click on a row to listen to the audio and read the transcript.
Loading corpus data…
0 matching records.
This is a sample of available metadata. For full metadata access, please request access.