3 Data

For this workshop, we use F0 traces from 220 recordings of 11 Mandarin speakers (5 female, 6 male) uttering ma with four different tones and five repetitions per tone. The tone labels 1–4 correspond to tones 55, 24, 213, 51, respectively. These recordings are freely available at the Production and Perception of Linguistic Voice Quality Project.

The traces comprise F0 measurements taken every 10 ms and extracted with Voicesauce, using the STRAIGHT algorithm (Liu and Kewley-Port 2004). These traces were assembled into a single CSV file Mandarin_ma_f0.csv available from the data directory.

# Analysis of the f0 traces of Mandarin speakers uttering MA.
#
# Created by Julian Villegas
# University of Aizu, 2026

This file can be imported into R:

f0_df <- read.csv('data/Mandarin_ma_f0.csv', 
    stringsAsFactors = T, 
    header=T)

Tone is a numeric value, but it should be considered a factor:

f0_df$Tone <- as.factor(f0_df$Tone)

We can inspect the resulting data frame now:

summary(f0_df)
##     Speaker     Tone     repetition   seg_Start     seg_End      
##  M1     :1685   1:3652   a:2911     Min.   :10   Min.   : 418.0  
##  F3     :1617   2:3698   b:2878     1st Qu.:10   1st Qu.: 570.0  
##  F2     :1502   3:3727   c:2856     Median :10   Median : 666.0  
##  F1     :1362   4:3280   d:2832     Mean   :10   Mean   : 680.2  
##  F5     :1353            e:2880     3rd Qu.:10   3rd Qu.: 786.0  
##  F6     :1296                       Max.   :10   Max.   :1002.0  
##  (Other):5542                                                    
##       t_ms              F0        
##  Min.   :  10.0   Min.   : 45.91  
##  1st Qu.: 170.0   1st Qu.:120.80  
##  Median : 330.0   Median :170.13  
##  Mean   : 345.3   Mean   :170.95  
##  3rd Qu.: 500.0   3rd Qu.:218.80  
##  Max.   :1000.0   Max.   :402.24  
## 
head(f0_df)
##   Speaker Tone repetition seg_Start seg_End t_ms      F0
## 1      F1    1          a        10     754   10 277.314
## 2      F1    1          a        10     754   20 265.947
## 3      F1    1          a        10     754   30 265.032
## 4      F1    1          a        10     754   40 267.388
## 5      F1    1          a        10     754   50 263.511
## 6      F1    1          a        10     754   60 289.069
str(f0_df)
## 'data.frame':    14357 obs. of  7 variables:
##  $ Speaker   : Factor w/ 11 levels "F1","F2","F3",..: 1 1 1 1 1 1 1 1 1 1 ...
##  $ Tone      : Factor w/ 4 levels "1","2","3","4": 1 1 1 1 1 1 1 1 1 1 ...
##  $ repetition: Factor w/ 5 levels "a","b","c","d",..: 1 1 1 1 1 1 1 1 1 1 ...
##  $ seg_Start : num  10 10 10 10 10 10 10 10 10 10 ...
##  $ seg_End   : num  754 754 754 754 754 754 754 754 754 754 ...
##  $ t_ms      : num  10 20 30 40 50 60 70 80 90 100 ...
##  $ F0        : num  277 266 265 267 264 ...

To ease the data analysis, we also want to add factors for

  • Gender
  • Audio file (here called token)
  • Where the trace start within the file (start_event)
f0_df$Gender <- factor(substr(f0_df$Speaker, 1, 1))
f0_df$token <- factor(paste(f0_df$Speaker, f0_df$Tone, 
            f0_df$repetition, sep = "_"))
f0_df$start_event <- !duplicated(f0_df$token)  

References

Liu, C., and D. Kewley-Port. 2004. “STRAIGHT: A New Speech Synthesiser for Vowel Formant Discrimination.” Acoustics Research Letters Online 5: 31–36.