14 Limitations

This is a very limited introduction to GAMs, as previously explained. The interested reader is welcome to explore these other topics that we could not cover for time reasons

  • Tensor product smooths (te()/ti()) for modeling genuine interactions between two continuous predictors, beyond the factor-by-smooth approach used here.

  • Non-Gaussian families. Our f0-in-semitones outcome was continuous and roughly normal, but mgcv supports binomial (e.g., presence/absence of a phonetic feature over time), Poisson (counts), and other families through the same machinery.

  • Automating data-quality checks. We manually flagged pitch-tracking errors by eye and by threshold today; in a real project, this detection is worth building into a pipeline that runs before any modeling — along with reconsidering Praat pitch floor/ceiling settings for tone categories (like our Tone 3) that are prone to creaky voice.

  • Principled model comparison. We compared a couple of candidate models with compareML(); with many candidate terms, prefer a theory-driven approach to model-building over comparing many models post hoc, and be mindful of the multiple-comparisons problem this creates.

  • Bayesian alternatives, such as brms or Bayesian workflows built on mgcv, offer another way to fit and interpret these models, with some advantages for uncertainty quantification.

  • Reporting standards for phonetics. For guidance on writing up GAMM results for a linguistics audience, see Sóskuthy (2017) and Wieling (2018) — both are approachable tutorial papers written specifically for our field.