Bayesian Inference Notes

1. Bayes Theorem

1.1. Bayes theorem

𝑝(πœƒ|𝐷)=𝑝(𝐷|πœƒ)𝑝(πœƒ)𝑝(𝐷)βˆπ‘(𝐷|πœƒ)𝑝(πœƒ)

1.2. Prior predictive distribution

𝑝(𝐷)=βˆ«Ξ˜π‘(𝐷|πœƒ)𝑝(πœƒ)dπœƒ

1.3. Posterior predictive distribution

𝑝(𝑦̃|𝐷)=βˆ«Ξ˜π‘(𝑦̃,πœƒ|𝐷)dπœƒ=βˆ«Ξ˜π‘(𝑦̃|πœƒ,𝐷)𝑝(πœƒ|𝐷)dπœƒ=βˆ«Ξ˜π‘(𝑦̃|πœƒ)𝑝(πœƒ|𝐷)dπœƒ

2. Fundamental Distributions

NamePDF/PMFMeanVarianceMode
Beta(𝑦|𝛼,𝛽)Ξ“(𝛼+𝛽)Ξ“(𝛼)Ξ“(𝛽)π‘¦π›Όβˆ’1(1βˆ’π‘¦)π›½βˆ’1𝛼𝛼+𝛽𝛼𝛽(𝛼+𝛽)2(𝛼+𝛽+1)π›Όβˆ’1𝛼+π›½βˆ’2
Binomial(𝑦|𝑛,𝑝)(𝑛𝑦)𝑝𝑦(1βˆ’π‘)π‘›βˆ’π‘¦π‘›π‘π‘›π‘(1βˆ’π‘)
Dirichlet(𝑦⃗|𝛼⃗)Ξ“(𝛼0)βˆπ‘–=1𝐾Γ(𝛼𝑖)βˆπ‘–=1πΎπ‘¦π‘–π›Όπ‘–βˆ’1𝛼𝑖𝛼0𝛼𝑖(𝛼0βˆ’π›Όπ‘–)𝛼02(𝛼0+1)π›Όπ‘–βˆ’1𝛼0βˆ’πΎ,𝛼𝑖>1βˆ€π‘–
Exponential(𝑦|πœ†)πœ†π‘’βˆ’πœ†π‘¦1πœ†ln2πœ†0
Erlang(𝑦|πœ†,π‘˜)πœ†π‘˜π‘¦π‘˜βˆ’1π‘’βˆ’πœ†π‘¦(π‘˜βˆ’1)!π‘˜πœ†π‘˜πœ†21πœ†(π‘˜βˆ’1)
ExGauss(𝑦|πœ‡,𝜎,πœ†)πœ†2exp(πœ†2(2πœ‡+πœ†πœŽ2βˆ’2𝑦))erfc(πœ‡+πœ†πœŽ2βˆ’π‘¦2𝜎
Gamma(𝑦|𝛼,𝛽)𝛽𝛼Γ(𝛼)π‘¦π›Όβˆ’1π‘’βˆ’π›½π‘¦π›Όπ›½π›Όπ›½2π›Όβˆ’1𝛽
InvGamma(𝑦|𝛼,𝛽)𝛽𝛼Γ(𝛼)π‘¦βˆ’π›Όβˆ’1π‘’βˆ’π›½/π‘¦π›½π›Όβˆ’1𝛽2(π›Όβˆ’1)2(π›Όβˆ’2)π›½π›Όβˆ’1
LogNormal(𝑦|𝛼,𝛽)1π‘¦πœŽ2πœ‹π‘’βˆ’(lnπ‘¦βˆ’πœ‡)22𝜎2π‘’πœ‡+𝜎22(π‘’πœŽ2βˆ’1)𝑒2πœ‡+𝜎2π‘’πœ‡βˆ’πœŽ2
Possion(𝑦|πœ†)πœ†π‘¦π‘’βˆ’πœ†π‘¦!πœ†πœ†
NegBinomial(π‘˜|π‘Ÿ,𝑝)(π‘˜+π‘Ÿβˆ’1π‘˜)(1βˆ’π‘)π‘˜π‘π‘Ÿπ‘Ÿ(1βˆ’π‘)π‘π‘Ÿ(1βˆ’π‘)𝑝2
Normal(𝑦|πœ‡,𝜎2)12πœ‹πœŽexp(βˆ’(π‘¦βˆ’πœ‡)22𝜎2)πœ‡πœŽ2πœ‡
Student(𝑦|𝜈)Ξ“(𝜈+12)πœ‹πœˆΞ“(𝜈2)(1+𝑦2𝜈)βˆ’πœˆ+120πœˆπœˆβˆ’20
Uniform(𝑦|π‘Ž,𝑏)1π‘βˆ’π‘Žπ‘Ž+𝑏2(π‘βˆ’π‘Ž)212
TableΒ 1: Common Probability Distributions

3. Functions

3.1. Beta Function

𝐡(𝑧1,𝑧2)=∫01𝑑𝑧1βˆ’1(1βˆ’π‘‘)𝑧2βˆ’1d𝑑

Properties:

4. Conjugate Prior

The idea of a conjugate prior is that for a given likelihood we choose a prior distribution such that, after observing data and applying Bayes’ theorem, the posterior distribution belongs to the same family as the prior.

That is, if 𝑝(πœƒ) and 𝑝(πœƒ|𝐷) have the same distributional form, then the prior is called a conjugate prior for the likelihood model.

This is useful because it makes Bayesian updating analytically tractable. Instead of performing difficult integration or numerical approximation, we can often derive the posterior parameters in closed form.

5. Conjugate Prior for Exponential Families

Note general exponential family:

𝑝(𝑦𝑖|πœƒ)=𝑓(𝑦𝑖)exp(πœ™(πœƒ)𝖳𝑒(𝑦𝑖)βˆ’π‘”(πœƒ))

For the data set 𝐷={𝑦1,…,𝑦𝑛} of i.i.d. observations, the likelihood is

𝑝(𝐷|πœƒ)=βˆπ‘–=1𝑛(𝑓(𝑦𝑖))exp(πœ™(πœƒ)π–³βˆ‘π‘–=1𝑛𝑒(𝑦𝑖)βˆ’π‘›π‘”(πœƒ))

So conjugate prior for that likelihood is

𝑝(πœƒ)∝exp(πœ™(πœƒ)π–³πœˆβˆ’π‘›0𝑔(πœƒ))

Posterior is

𝑝(πœƒ|𝐷)∝exp(πœ™(πœƒ)𝖳(𝜈+βˆ‘π‘–=1𝑛𝑒(𝑦𝑖))βˆ’(𝑛0+𝑛)𝑔(πœƒ))

6. Proper and Improper Prior Distributions

A prior is called proper if it is a valid probability distribution:

𝑝(πœƒ)β‰₯0,βˆ€πœƒβˆˆΞ˜,βˆ«πœƒβˆˆΞ˜π‘(πœƒ)dπœƒ=1

And improper if

𝑝(πœƒ)β‰₯0,βˆ€πœƒβˆˆΞ˜,βˆ«πœƒβˆˆΞ˜π‘(πœƒ)dπœƒ=∞

In theory, all priors are acceptable, as long as the posterior is proper.

7. Fisher Information Matrix

𝐼⃗(πœƒβƒ—)=𝐸((𝛁log𝑝(𝑦|πœƒβƒ—))(𝛁log𝑝(𝑦|πœƒβƒ—))𝖳)=βˆ’πΈ(𝛁2log𝑝(𝑦|πœƒβƒ—))

8. Jeffreys’ Prior

πœ‹π½(πœƒβƒ—)∼|𝐼⃗(πœƒβƒ—)|2

9. Pivotal Quantities

For the binomial and other single-parameter models, different principles give (slightly) different noninformative prior distributions. But for two casesβ€”location parameters and scale parametersβ€”all principles seem to agree[1].

9.1. Location Parameter

𝑝(πœƒ)∼1

9.2. Scale Parameter

𝑝(πœƒ)∼1πœƒ

10. Predictive Accuracy

People care about the accuracy in two different ways. First to assume that the model is all we known and check posterior predictions. The second is to compare several candidate models. Even if all of the models being considered have mismatches with the data, it can be informative to evaluate their predictive accuracy, compare them, and consider where to go next[2].

11. KL Divergence

12. Linear Algebra

12.1. Convex Combination

A subset 𝐴 of a vector space 𝑉 is said to be convex if πœ†π‘₯βƒ—+(1βˆ’πœ†)𝑦⃗ for all vectors π‘₯βƒ—,π‘¦βƒ—βˆˆπ΄, and all scalars πœ† in [0,1].

Via induction, this can be seen to be equivalent to the requirement that πœ†π‘₯βƒ—1,πœ†π‘₯βƒ—2,…,πœ†π‘₯βƒ—π‘›βˆˆπ΄ for all vectors π‘₯βƒ—1,π‘₯βƒ—2,…,π‘₯βƒ—π‘›βˆˆπ΄, and for all scalars πœ†1,πœ†2,…,πœ†π‘›β‰₯0 such that βˆ‘π‘˜π‘–=1.

13. Reversible Jump MCMC

Reversible Jump Markov chain Monte Carlo (RJMCMC) samples jointly over a family of candidate models β„³οΈ€={π‘€π‘˜:π‘˜βˆˆπΎ} and their parameters. Under model π‘€π‘˜, let πœƒβƒ—π‘˜βˆˆΞ˜π‘˜βŠ†β„π‘›π‘˜. For a fixed observed data set 𝐷, the joint posterior is the product of the likelihood and the joint prior, 𝑝(π‘˜,πœƒβƒ—π‘˜)=𝑝(πœƒβƒ—π‘˜|π‘˜)𝑝(π‘˜)[3]:

πœ‹(π‘˜,πœƒβƒ—π‘˜|𝐷)=β„’οΈ€(𝐷|π‘˜,πœƒβƒ—π‘˜)𝑝(πœƒβƒ—π‘˜|π‘˜)𝑝(π‘˜)βˆ‘π‘šβˆˆπΎβˆ«Ξ˜π‘šβ„’οΈ€(𝐷|π‘š,πœƒβƒ—π‘š)𝑝(πœƒβƒ—π‘š|π‘š)𝑝(π‘š)dπœƒβƒ—π‘š.

Unlike ordinary MCMC, the state space is the disjoint union βŠπ‘˜βˆˆπΎ({π‘˜}Γ—Ξ˜π‘˜), so a jump between π‘€π‘˜ and π‘€π‘˜β€² may change the dimension of the parameter vector.

13.1. Sampler Configuration

An RJMCMC sampler is specified by the following components:

Starting from (π‘˜,πœƒβƒ—π‘˜), select a destination model π‘˜β€² with probability π‘Ÿ(π‘˜β€²|π‘˜), draw π‘’βƒ—βˆΌπ‘žπ‘˜,π‘˜β€²(β‹…|πœƒβƒ—π‘˜), and apply

(πœƒβƒ—π‘˜β€²,𝑒⃗′)=π‘‡π‘˜,π‘˜β€²(πœƒβƒ—π‘˜,𝑒⃗)

The corresponding Jacobian factor is:

π½π‘˜,π‘˜β€²(πœƒβƒ—π‘˜,𝑒⃗)=|det(πœ•πœƒβƒ—π‘˜β€²,π‘’βƒ—β€²πœ•(πœƒβƒ—π‘˜,𝑒⃗))|

Accept the proposed state (π‘˜β€²,πœƒβƒ—π‘˜β€²) with probability:

𝛼((π‘˜,πœƒβƒ—π‘˜),(π‘˜β€²,πœƒβƒ—π‘˜β€²))=min(1,πœ‹(π‘˜β€²,πœƒβƒ—π‘˜β€²|𝐷)π‘Ÿ(π‘˜|π‘˜β€²)π‘žπ‘˜β€²,π‘˜(𝑒⃗′|πœƒβƒ—π‘˜β€²)πœ‹(π‘˜,πœƒβƒ—π‘˜|𝐷)π‘Ÿ(π‘˜β€²|π‘˜)π‘žπ‘˜,π‘˜β€²(𝑒⃗|πœƒβƒ—π‘˜)π½π‘˜,π‘˜β€²(πœƒβƒ—π‘˜,𝑒⃗))

The posterior’s normalizing constant cancels from this ratio. If the proposal is rejected, retain (π‘˜,πœƒβƒ—π‘˜). With the remaining probability 1βˆ’βˆ‘π‘˜β€²β‰ π‘˜π‘Ÿ(π‘˜β€²|π‘˜), apply πΎπ‘˜ to update the parameters within the current model. Together, these updates preserve πœ‹(π‘˜,πœƒβƒ—π‘˜|𝐷) as the stationary distribution.

Bibliography

  • [1] A. Gelman, J. B. Carlin, H. S. Stern, D. B. Dunson, A. Vehtari, and others, Bayesian Data Analysis, Third (Crc, Boca Raton, Florida, 2013).
  • [2] A. Gelman, J. Hwang, and A. Vehtari, Understanding predictive information criteria for Bayesian models, Statistics and Computing 24, 997 (2014), https://doi.org/10.1007/s11222-013-9416-2.
  • [3] S. Brooks, A. Gelman, G. Jones, and X.-L. Meng, Handbook of Markov Chain Monte Carlo, 1st edition (Chapman, Hall/CRC, Boca Raton London New York, 2011).