I've been making an effort to cover extra reading material in my free time. I thought this course would provide an alternative perspective on statistical thinking, especially not having taken a formal stats course (and also having a deep distrust of hypothesis testing).
I've finally finished the homework for the first week, so here are (I think) the main ideas:
The data comes from somewhere: The data is generated through a process. Perhaps it is drawn from a uniform distribution, and the grad student experimenter has a small chance of writing down the wrong values. Or, perhaps it depends on the measured position of the Moon and Mars, both with experimental error. By using Directed Acyclic Graphs (DAG) we represent our conceptual prior of how variables affect each other to generate the observations.
The tree in the forest: We typically use best-fit points or Maximum Likelihood Estimates to describe the model and to make predictions. Bayesian rethinking tells us that the model parameters have a probability distribution. By using that as a weight and integrating that with the data generating process, we obtain the average of all possible realisations in the posterior predictive distribution.
Prior beliefs are everywhere: The choice of model, or DAG, or initial parameter distribution is a prior that we should state explicitly - it is better to see our assumptions clearly from the get go. The belief that our data follows some distribution in asymptotic limits (if that even exists) is a sort of prior too.
I'm enjoying this course. It stresses the fact that statistical thinking is a golem - a powerful machine that is possibly destructive if misunderstood, and provides tools (new to me) that help to clarify thinking. I'm looking forward to the rest of the lectures.
No comments:
Post a Comment