A maximum entropy principle selects, when a maximizer exists, a PP^\star from a nonempty feasible class C\mathcal C such that

Parg maxPCH(P),P^\star \in \operatorname*{arg\,max}_{P\in\mathcal C} H(P),

where HH is the chosen entropy functional and C\mathcal C is the feasible class of distributions.

The guiding idea is to choose a distribution that adds as little structure as the entropy model permits beyond the stated constraints. This depends on the chosen entropy and, in the continuous case, on the underlying coordinates or . Relative-entropy minimization against a specified reference distribution is the corresponding reference-dependent formulation.

Examples

For discrete laws, HH is commonly ; for absolutely continuous laws it is often . A typical feasible class is specified by support or constraints, for example EP[gi(X)]=ci\mathbb E_P[g_i(X)]=c_i.

  • On a finite set of nn outcomes, the maximizes Shannon entropy.
  • Among probability distributions on R\mathbb R with a fixed mean and a fixed positive , the with those parameters maximizes differential entropy.