In probability theory and statistics, the continuous binomial distribution (also called the cobin distribution) is a family of continuous probability distributions on the unit interval that belongs to an exponential dispersion family. It was introduced as a response distribution for generalized linear models for continuous proportional data, proposed as an alternative to beta regression. The special case λ = 1 {\displaystyle \lambda =1} coincides with the continuous Bernoulli distribution.

Definition

A random variable X {\displaystyle X} is said to follow a continuous binomial (cobin) distribution with natural parameter θ {\displaystyle \theta } and inverse dispersion λ ∈ { 1 , 2 , … } {\displaystyle \lambda \in \{1,2,\dots \}}, written X ∼ c o b i n ( θ , λ − 1 ) {\displaystyle X\sim \mathrm {cobin} (\theta ,\lambda ^{-1})}, if it has density on [ 0 , 1 ] {\displaystyle [0,1]} given by

f ( x ; θ , λ ) = h ( x ; λ ) exp ( λ θ x − λ B ( θ ) ) , 0 ≤ x ≤ 1 , {\displaystyle f(x;\theta ,\lambda )=h(x;\lambda )\exp \!{\big (}\lambda \theta x-\lambda B(\theta ){\big )},\qquad 0\leq x\leq 1,}

where the log-partition function is

B ( θ ) = { log ( e θ − 1 θ ) , θ ≠ 0 , 0 , θ = 0 , {\displaystyle B(\theta )={\begin{cases}\log \!\left({\frac {e^{\theta }-1}{\theta }}\right),&\theta \neq 0,\\0,&\theta =0,\end{cases}}}

and the base measure h ( x ; λ ) {\displaystyle h(x;\lambda )} is

h ( x ; λ ) = λ ( λ − 1 ) ! ∑ k = 0 λ ( − 1 ) k ( λ k ) max ( 0 , λ x − k ) λ − 1 , {\displaystyle h(x;\lambda )={\frac {\lambda }{(\lambda -1)!}}\sum _{k=0}^{\lambda }(-1)^{k}{\lambda \choose k}\,\max(0,\lambda x-k)^{\lambda -1},}

with h ( x ; 1 ) = 1 {\displaystyle h(x;1)=1}. The function h ( x ; λ ) / λ {\displaystyle h(x;\lambda )/\lambda } coincides with the probability density function of the Irwin–Hall distribution with parameter n = λ {\displaystyle n=\lambda }, evaluated at λ x {\displaystyle \lambda x}.

When λ {\displaystyle \lambda } is fixed, the cobin distribution belongs to a one-parameter natural exponential family in θ {\displaystyle \theta }.

Related distributions

  • Bates distribution: when θ = 0 {\displaystyle \theta =0}, the density reduces to h ( x ; λ ) {\displaystyle h(x;\lambda )}, corresponding to the distribution of the mean of λ {\displaystyle \lambda } independent U n i f o r m ( 0 , 1 ) {\displaystyle \mathrm {Uniform} (0,1)} random variables (equivalently, a scaled Irwin–Hall distribution or Bates distribution).
  • Uniform distribution: when λ = 1 {\displaystyle \lambda =1} and θ = 0 {\displaystyle \theta =0}, the distribution reduces to the continuous uniform distribution on [ 0 , 1 ] {\displaystyle [0,1]}.
  • If X 1 , … , X λ {\displaystyle X_{1},\dots ,X_{\lambda }} are independent and identically distributed continuous Bernoulli random variables with common natural parameter θ {\displaystyle \theta }, then

X ¯ = 1 λ ∑ i = 1 λ X i ∼ c o b i n ( θ , λ − 1 ) . {\displaystyle {\bar {X}}={\frac {1}{\lambda }}\sum _{i=1}^{\lambda }X_{i}\sim \mathrm {cobin} (\theta ,\lambda ^{-1}).}

Properties

Mean and variance

The mean and variance of X ∼ c o b i n ( θ , λ − 1 ) {\displaystyle X\sim \mathrm {cobin} (\theta ,\lambda ^{-1})} can be expressed in terms of derivatives of B ( θ ) {\displaystyle B(\theta )}:

  • E ⁡ ( X ) = B ′ ( θ ) = e θ e θ − 1 − 1 θ {\displaystyle \operatorname {E} (X)=B'(\theta )={\frac {e^{\theta }}{e^{\theta }-1}}-{\frac {1}{\theta }}}, for θ ≠ 0 {\displaystyle \theta \neq 0}.
  • Var ⁡ ( X ) = 1 λ B ″ ( θ ) = 1 λ ( 1 θ 2 − e θ ( e θ − 1 ) 2 ) {\displaystyle \operatorname {Var} (X)={\frac {1}{\lambda }}B''(\theta )={\frac {1}{\lambda }}\left({\frac {1}{\theta ^{2}}}-{\frac {e^{\theta }}{(e^{\theta }-1)^{2}}}\right)}, for θ ≠ 0 {\displaystyle \theta \neq 0}.

If θ = 0 {\displaystyle \theta =0}, then E ⁡ ( X ) = 1 2 {\displaystyle \operatorname {E} (X)={\frac {1}{2}}} and Var ⁡ ( X ) = 1 12 λ {\displaystyle \operatorname {Var} (X)={\frac {1}{12\lambda }}}.

Sufficient statistic for the mean

If X 1 , … , X n {\displaystyle X_{1},\ldots ,X_{n}} are independent and identically distributed continuous binomial random variables with common natural parameter θ {\displaystyle \theta } and fixed inverse dispersion parameter λ {\displaystyle \lambda }, then the sample mean

X ¯ = 1 n ∑ i = 1 n X i {\displaystyle {\bar {X}}={\frac {1}{n}}\sum _{i=1}^{n}X_{i}}

is a sufficient statistic for θ {\displaystyle \theta }.

This is in contrast with the beta distribution: under a mean–precision parameterisation X i ∼ B e t a ( μ ϕ , ( 1 − μ ) ϕ ) {\displaystyle X_{i}\sim \mathrm {Beta} (\mu \phi ,(1-\mu )\phi )} with fixed ϕ {\displaystyle \phi }, a sufficient statistic for the mean μ {\displaystyle \mu } is

∑ i = 1 n log ( X i 1 − X i ) , {\displaystyle \sum _{i=1}^{n}\log \!\left({\frac {X_{i}}{1-X_{i}}}\right),}

not the sample mean X ¯ {\displaystyle {\bar {X}}}.

Applications

The cobin distribution has been proposed as a response distribution for generalized linear models of continuous proportional data, as an alternative to beta regression, including extensions with random effects.