Basic root structure in Proto-Uralic
Before I go on, time for a small general overview I’ve been meaning to write for a while, to keep my readership on board.
Word structure in the languages of the world varies greatly. In most recorded or reconstructed languages of the world, not all permissible word shapes are however equally common: certain phonotactic shapes could be considered “typical”, and others “marginal”. Proto-Uralic is no exception to this.
According to long-standing general consensus (the reasons for which would not be difficult to demonstrate, but this would take too long here), most typically PU roots were made of two syllables. These were not equivalent however: the second, unstressed syllable was almost always a simple short, open syllable with a single initial consonant (though word-final consonants could occur in inflection), while the initial, stressed syllable could well have a final consonant, or lack an initial one. So a generalized syllable structure cannot really be stated: “(C)V(C)” would be overly general, “CV” would be too limited. We can easily join these for a “root structure” formula though: (C)V(C)CV.
The asymmetry between stressed and unstressed syllables goes deeper still though. All reconstructions agree that at least six (but probably at least eight) vowels were distinguished in the initial syllable — and that no more than two options were consistently available in the next one. It is likely (but not universally accepted) that in part this was due to vowel harmony: the corresponding back and front open vowels *a and *ä were not contrasted, but both words such as *pala “bit” and *pälä “half” could still occur, ie. with *a found following back vowels, and *ä following front vowels. [1]
Aside from the open vowel option, another quality also occurred. What specific sound value(s) this had is a disputed issue. Maximal phonetic differentiation would call for a close vowel such as *i (possibly with a vowel-harmonic counterpart *ï), and this is the transcription that may have the widest use in the literature these days: e.g. *weti “water”. However, the actual evidence does not favor this: the non-open vowel is widely across Uralic rather reflected as a reduced vowel, something like *ə. Even Finnic /i/ only occurs word-finally, while in other positions /e/ is found (which adds up to alternation in e-stem nouns). [2]
The current transcription I use on this blog follows the scheme a/ä/ə. For an original open stem vowel, I write *a after back vowels and *ä after front vowels. If I need to speak of these classes as a whole I use the cover symbol *A. For an original non-open stem vowel, I write *ə. Therefore, a reconstruction such as *wetə is completely equivalent with standard *weti.
You may occasionally also see the symbol ɜ. This is traditional Uralistic transcription for a vowel whose quality cannot be determined. Certain Uralic languages have lost or merged all unstressed vowels, and today their basic inherited roots have a monosyllabic shape, most commonly (C)V(C)(C). For PU roots that have not survived in one of the three branches that fairly consistently preserve unstressed vowels (Finnic, Samic, and Samoyedic), the contrast between *A and *ə may not be recoverable. I find this less intrusiv than the practice common elsewhere in historical linguistics of simply writing “V” for “vowel”: making up an example, compare *tirɜ vs. *tirV?
(Another reconstruction notation worth mentioning at this point: while *asterisks normally mark reconstructed items, I’ve picked up from somewhere the idea to use #hashes for approximate pseudo-reconstructions that don’t actually rely on known regular correspondences. This is useful with e.g. “messy” roots that show very divergent shapes all over the family, or when referring to data from families I am not particularly familiar on.)
In addition to *A-stems and *ə-stems, a third type has also been considered fairly basic. Some Finnic roots show what are called “primary” long vowels, corresponding to single vowels in most other Uralic languages. [3] A good example may be the root for “language”: Finnish kieli (kiele-), Northern Sami giella, Mokša /käĺ/, Udmurt /kɨl/, Komi /kɨv/, Khanty *kööɬ, Nganasan /kieja/, to pick a few reflexes. Curiously these only occurred with the stem vowel *-ə. This led to a third basic root type with assumed original long vowels, *(C)VVCə (or *(C)VVCi) hanging around in reconstructions for a while.
A full coverage on the story of these would be too much a digression (and I’ve covered some of it before), but suffice to say that these days it seems that long vowels, or anything equivalent, is actually not necessary for Proto-Uralic after all. The “language” root above may be reconstructed as *kälə, for example. I point those interested in the details towards a recent article by A. Aikio. [3]
Many further observations on PU root structure would be possible, but this should suffice to keep readers on board with some future posts.
[1] The International Phonetic Alphabet defines [a] as an open front or central vowel, but in Uralistics, *a and *ä refer to the vowels known as [ɑ] and [æ] in the IPA, after the orthography of Finnish and Estonian.
[2] A recent article exporing this idea is Petri Kallio (2012): The non-initial-syllable vowel reductions from Proto-Uralic to Proto-Finnic (SUST 264). Some of the topics are better explained by Aikio in the same volume (see note 4 below), but I find the basic thesis on *ə being a better reconstruction than *i convincing.
[3] This is in contrast to long vowels that have arisen from loss of former consonants and thus correspond to VC(V) structures in (some) other Uralic languages.
[4] Ante Aikio (2012): On Finnic long vowels, Samoyed vowel sequences, and Proto-Uralic *x (SUST 264)