Hey nostalgebraist, I have a question. (Asking here and not in the ask box because (1) my followers might be interested, and (2) it's long. Hope you don't mind, and obviously, feel free not to answer!)
I'm a grad student in... well, I don't know what field I'm in, but let's say computational cognitive science. In any case, I'm especially interested in probabilistic models of concept learning, and most of these seem to be Bayesian. Talking to other researchers in this field, I gather that people think priors come from the following sources: (1) evolution, so you're born with your priors, (2) some kind of hierarchical modeling thing, where e.g. if you learn a new concept like goat, your prior on the mean/variance will come from your experience with other, similar concepts, like sheep and horse, or (3) sequential Bayesian updates, so your prior at time t comes from your posterior at time t-1.
In practice, when people are building models, they often pick the prior which works best empirically (that is, people tune the hyperparameters on held-out data); this seems philosophically justifiable, since evolution is presumably also experimenting and selecting the priors which perform best in practice. Also, as you mentioned, people often fold Occam's Razor into the prior, which seems to work empirically. Though I don't necessarily believe that "the simplest explanation is most likely to be true"; I think Occam's Razor works well because (1) complicated explanations make it easier to overfit, and (2) we as humans have finite computational resources and simpler explanations are easier to use in reasoning.
Anyway, none of this seems objectionable to me. And when I've read LW, it seemed like people there were treating priors in the same way. Have I been misreading LW/subconsciously steelmanning LW to the above? Or is there something here that you find objectionable?
(For the record, I have a heaping pile of problems with LW's use of Bayesianism. It's just that priors isn't one of them, so I was curious whether I was missing something. My main problem is that LW treats Bayesianism as normative rather than descriptive. Like, people are supposed to figure out their probabilities and do actual calculations with them, which is just, I can't even. I mean (1) you have little hope of estimating your subjective probabilities accurately, and (2) this is like the least efficient form of reasoning ever. Academic Bayesianism is much more sensible; in the computational cog sci literature, there's a big debate between people who say "humans already behave subconsciously like Bayesian reasoners" and people who say "humans behave suboptimally because human reasoning is full of biases and heuristics", but neither of these groups says "and humans would be better at reasoning if they did explicit, conscious calculations with 'probabilities' that they pulled out of their ass".)