So I did a thing...
I took all the words found at this place: http://www.mieliestronk.com/corncob_lowercase.txt
Which surmounts to about 58110 words.
Then I coded a small piece to get the letter fequencies and letter combinations. Mostly in order to create random words for my roleplaying group... I, err, may have gone a bit over board there.
Anyway!
I won’t show all the data here, cuss it’s 15.000 rows long, but I found some curious tidbits at least. Like the only double letter combination that ends words are ‘LZ‘ and ‘VS‘. There are others that almost always end words, of course, but they all can have at least another letter as well.
So let’s look at letters that start words!
-- first letter -- s: 0.1148 percent c: 0.0945 percent p: 0.0785 percent d: 0.065 percent r: 0.0625 percent a: 0.0599 percent b: 0.0552 percent m: 0.0507 percent t: 0.0496 percent i: 0.046 percent e: 0.0445 percent f: 0.044 percent h: 0.0349 percent u: 0.0331 percent l: 0.0316 percent g: 0.0316 percent w: 0.0265 percent o: 0.0239 percent n: 0.0158 percent v: 0.014 percent j: 0.0081 percent k: 0.0061 percent q: 0.005 percent y: 0.0025 percent z: 0.0015 percent x: 0.0002 percent
So yeah, the most common first letter is S with the least common X.
No love for X ;-;
So... what’s the point of all this?
None really... I did it because I wanted to be able to generate random words that actually are pronouncable. Which might’ve worked? I’ll let you be the judge of that!
Hled dess refenratings fies stroy taticellacch!
P.S If anyone wishes to know more about this, or is interested in a 227kb large file with the analysis, just ask me and I’ll be happy to talk more about it.











