Appendix B — Probability and Stochastic Processes Review

NoteLearning objectives

After reading this chapter you should be able to:

  • compute expectations and variances, and condition on a second random variable
  • find the mean and the variance of a sum, and say when the variances add
  • state what the central limit theorem gives and what it does not
  • work with the Poisson and renewal processes used in the inventory models
  • explain why the fill rate and the ready rate coincide only under Poisson demand
  • say when the variance of a duration matters to a total and when only its mean does
  • compute the expectations of random sums that arise in lead-time demand

B.1 Random Variables and Expectation

This section and the four that follow are a review for a reader who has had a probability course, not a substitute for one. They collect the results the inventory models use, in the notation the book uses, so that a chapter can point here rather than rederive them.

A random variable assigns a number to the outcome of an experiment whose result is not known in advance. Let \(X\) represent one. The cumulative distribution function

\[ F(x) = P\{X \le x\} \tag{B.1}\]

determines everything about \(X\), and it is the object the inventory models actually use. Recall from Equation 7.12 that the newsvendor quantity is a quantile of \(F\), and from Section C.3 that both loss functions are integrals of it.

A discrete random variable takes values in a countable set and carries a probability mass function \(g(x) = P\{X = x\}\). A continuous one carries a density \(f(x)\) with \(F(x) = \int_{-\infty}^{x}f(u)\,du\). Demand arrives in whole units, so the natural default is discrete and a continuous model is an approximation made for convenience. Section 7.8 says when that approximation is safe and what it costs when it is not.

B.1.1 Expectation, Variance, and Two Ratios

The expectation of \(X\) is its probability-weighted average,

\[ E[X] = \sum_{x} x\,g(x) \quad\text{or}\quad \int_{-\infty}^{\infty} x f(x)\,dx \tag{B.2}\]

and the variance is the expected squared deviation from it,

\[ \mathit{Var}[X] = E\big[(X - E[X])^{2}\big] = E[X^{2}] - \big(E[X]\big)^{2} \tag{B.3}\]

The second form is the one to compute with, and it is how Equation B.31 below collapses to a second moment. The standard deviation \(\sigma_{X} = \sqrt{\mathit{Var}[X]}\) carries the units of \(X\) and is therefore the one to report.

Two dimensionless ratios appear throughout the book and they are not the same thing. The coefficient of variation \(c_{X} = \sigma_{X}/E[X]\) measures spread relative to level, and Table C.1 reads a normal model’s defect off it directly. The variance to mean ratio of Equation C.1 selects among discrete families and carries the units of \(X\). Notice that \(c_{X}\) compares a standard deviation with a mean and \(\mathit{VMR}\) compares a variance with a mean, so the two answer different questions and only one of them is a pure number.

For a linear transformation, with \(a\) and \(b\) constants,

\[ E[aX + b] = aE[X] + b, \qquad \mathit{Var}[aX + b] = a^{2}\mathit{Var}[X] \tag{B.4}\]

Notice that \(b\) moves the mean and does not touch the variance, and that \(a\) enters the variance squared. That is why a change of units rescales a standard deviation linearly, and it is the whole of the argument in Section 1.4.1 that a constant added to every alternative cannot change which alternative is least.

B.1.2 A Function of a Random Variable

For any function \(h\),

\[ E\big[h(X)\big] = \sum_{x} h(x)\,g(x) \quad\text{or}\quad \int_{-\infty}^{\infty} h(x) f(x)\,dx \tag{B.5}\]

so the expectation is taken against the distribution of \(X\) and no distribution of \(h(X)\) has to be found. Equation 7.10 is Equation B.5 applied to the newsvendor cost, and Equation C.9 is it applied to \((X-b)^{+}\).

The result that matters here is the one about what you may not do. In general

\[ E\big[h(X)\big] \ne h\big(E[X]\big) \tag{B.6}\]

and the two sides differ by more than a rounding. Replacing a random demand by its average and evaluating the cost there is the commonest modeling error in inventory work, and it has a name outside this book: the flaw of averages. Thus, a model that plans against average demand is not a conservative approximation of a stochastic model. It is a different model, and usually an optimistic one.

For example, take the five-point demand sketch of Table 7.1, whose mean is 485 belts, with \(c_{o} = \$38\) and \(c_{u} = \$45\). Order the average, \(Q = 485\). Substituting the average demand into Equation 7.4 gives

\[ C(485, 485) = 38(485-485)^{+} + 45(485-485)^{+} = \$0 \]

which reports that ordering the mean costs nothing at all. Now let’s take the expectation properly. We compute the expected leftover as 49.0 belts and the expected shortage as also 49.0 belts, so by Equation 7.10

\[ E\big[C(485)\big] = 38(49.0) + 45(49.0) = 49.0(83) = \$4{,}067 \]

The flaw of averages understated the cost of that decision by the entire cost of it. Notice that the two expectations came out equal, and that this is not a coincidence: Equation 7.9 says their difference is \(Q - E[X]\), which is zero at \(Q = E[X]\). That identity is worth using as an arithmetic check every time a loss function is evaluated at the mean.

B.1.3 Conditioning

We derive most of the results in this book by conditioning on a second random variable and then averaging over it. Let \(Y\) represent that second variable. The conditional expectation \(E[X \mid Y]\) is the expectation of \(X\) computed within each value of \(Y\), and it is itself a random variable, because it depends on which value \(Y\) takes. Averaging it over \(Y\) recovers the unconditional mean,

\[ E[X] = E\big[E[X \mid Y]\big] \tag{B.7}\]

which is the law of total expectation. Equation B.27 is Equation B.7 with \(Y\) the number of terms in a sum, and Chapter 8 conditions on the inventory position in the same way.

The variance needs both a within and a between term. The law of total variance is

\[ \mathit{Var}[X] = E\big[\mathit{Var}[X \mid Y]\big] + \mathit{Var}\big[E[X \mid Y]\big] \tag{B.8}\]

That is, the total variability is the average variability within a value of \(Y\) plus the variability of the averages across values of \(Y\). Both terms are non-negative, so conditioning on anything can only reduce the first term at the price of raising the second. Equation B.29 is Equation B.8 worked through for a random sum, and the asymmetry that section makes so much of is visible here already: the first term averages a variance and the second takes the variance of a mean.

Almost every quantity the inventory models need is a sum of the quantities defined here, so the next section takes up what addition does to a mean, to a variance, and to a distribution.

B.2 Sums of Random Variables

Almost every quantity this book models is a sum. Demand over a lead time is a sum of period demands, demand over a season is a sum of weeks, and an aggregate investment is a sum across items. This section gives what happens to the moments and to the distribution when quantities are added.

Let \(X_{1}, X_{2}, \ldots, X_{n}\) represent random variables and let \(S_{n} = X_{1} + X_{2} + \cdots + X_{n}\).

B.2.1 The Mean and the Variance of a Sum

The mean of a sum is the sum of the means, always, whatever the dependence among the terms:

\[ E[S_{n}] = \sum_{i=1}^{n} E[X_{i}] \tag{B.9}\]

The variance is not so obliging. In general we have

\[ \mathit{Var}[S_{n}] = \sum_{i=1}^{n}\mathit{Var}[X_{i}] + 2\sum_{i<j}\mathit{Cov}[X_{i}, X_{j}] \tag{B.10}\]

and the covariance terms vanish only when the terms are uncorrelated. Thus, the familiar statement that variances add is a statement about independence and not about addition. Notice the direction of the error when independence is assumed wrongly. Demand that runs in streaks has positive covariances, so Equation B.10 without its second term understates the spread, and understating the spread understates the safety stock. Section B.5.5 makes the same point for random sums.

When the terms are independent and identically distributed, with mean \(\mu\) and variance \(\sigma^{2}\),

\[ E[S_{n}] = n\mu, \qquad \mathit{Var}[S_{n}] = n\sigma^{2}, \qquad \sigma_{S_{n}} = \sigma\sqrt{n} \tag{B.11}\]

Notice that the mean grows linearly and the standard deviation grows only as \(\sqrt{n}\). Thus, the coefficient of variation of a sum is

\[ c_{S_{n}} = \frac{\sigma\sqrt{n}}{n\mu} = \frac{c_{X}}{\sqrt{n}} \tag{B.12}\]

so aggregation makes a quantity relatively less variable even as it makes it absolutely more variable. That single fact is behind a great deal of inventory practice: it is why pooling stock across locations reduces the safety stock needed, which Chapter 9 takes up, and it is why a demand that is far too variable to model in one bucket can be tractable in a larger one.

For example, weekly demand for the belt of Section 7.3 has a mean of 33.2404 belts and a standard deviation of 24.7596 belts, so \(c_{X} = 0.7449\). Over the fourteen remaining weeks,

\[ E[X] = 14(33.2404) = 465.37 \text{ belts}, \qquad \mathit{Var}[X] = 14(613.0387) = 8582.54 \text{ belts}^{2} \]

so \(\sigma = 92.64\) belts and \(c_{X} = 92.64/465.37 = 0.1991\). Let’s check that against Equation B.12: \(0.7449/\sqrt{14} = 0.7449/3.7417 = 0.1991\). The two agree, and that check is worth performing every time, because it catches a variance that was added when it should have been multiplied.

B.2.2 Averages, and Why Precision Is Expensive

The sample mean of \(n\) independent observations is \(\bar{X} = S_{n}/n\). Applying Equation B.4 to Equation B.11 with \(a = 1/n\),

\[ E[\bar{X}] = \mu, \qquad \mathit{Var}[\bar{X}] = \frac{\sigma^{2}}{n}, \qquad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}} \tag{B.13}\]

The quantity \(\sigma_{\bar{X}}\) is the standard error of the mean, and an approximate confidence interval for \(\mu\) is \(\bar{X} \pm z\sigma/\sqrt{n}\), whose half-width is \(z\sigma/\sqrt{n}\).

Read the exponent before using it. The half-width falls with \(\sqrt{n}\), so halving it costs four times the observations and dividing it by ten costs a hundred times. Section 7.10 reports every simulated figure with its half-width for this reason, and the arithmetic above is why a simulated answer is expensive to sharpen and an exact one is not.

B.2.3 Families Closed Under Addition

Equation B.11 gives two moments and not a distribution. For a few families the distribution of the sum stays in the family, and those cases turn an approximation into an identity.

Table B.1: Families closed under addition. In every row the parameter that is shared must be shared, and the parameter that adds is the one that counts.
If \(X_{1}, \ldots, X_{n}\) are independent and then \(S_{n}\) is
normal, \(\mathit{N}(\mu_{i}, \sigma_{i}^{2})\) normal, \(\mathit{N}(\sum\mu_{i}, \sum\sigma_{i}^{2})\)
Poisson, rate \(\lambda_{i}\) Poisson, rate \(\sum\lambda_{i}\)
gamma with a common scale, shapes \(\alpha_{i}\) gamma, same scale, shape \(\sum\alpha_{i}\)
negative binomial with a common \(p\), sizes \(r_{i}\) negative binomial, same \(p\), size \(\sum r_{i}\)
binomial with a common \(p\), trials \(n_{i}\) binomial, same \(p\), trials \(\sum n_{i}\)

Three notes on Table B.1. The normal row holds without the terms being identically distributed, which is the reason the normal is so often reached for. The gamma and negative binomial rows require the common parameter, so a sum of gammas with different scales is not a gamma and has to be approximated. And the exponential is the gamma with shape one, so a sum of \(n\) independent exponentials with the same rate is the Erlang of Section C.2.2, which is the fact Chapter 8 needs when a lead time is built from stages.

Closure is the exception and not the rule. For every other combination the distribution of the sum has to be obtained by convolution, by a transform, by simulation, or by matching moments to a chosen family as Section C.4.2 does.

B.2.4 The Central Limit Theorem, and What It Does Not Say

If \(X_{1}, X_{2}, \ldots\) are independent and identically distributed with mean \(\mu\) and finite variance \(\sigma^{2}\), then

\[ \frac{S_{n} - n\mu}{\sigma\sqrt{n}} \;\longrightarrow\; \mathit{N}(0, 1) \qquad\text{as } n \to \infty \tag{B.14}\]

whatever the distribution of the terms. Thus, a quantity aggregated over many independent contributions tends toward the normal, which is the justification Section C.2.3 gives for a family that is otherwise indefensible for demand.

Three qualifications, and you should apply all three before treating an aggregate as normal.

It is a statement about the limit, not about your \(n\). How large \(n\) must be depends on how skewed the terms are. A symmetric term converges in a handful of observations and a heavily right-skewed one may need hundreds.

The approximation is worst in the tail, which is where inventory lives. Recall from Section 7.6 that the quantity of interest is a quantile, often far from the mean. A central limit argument is strongest at the center, so agreement of the means says nothing about agreement at the 95th percentile.

The support is wrong. The normal puts mass below zero and demand cannot be negative. Table C.1 prices that defect against \(c_{X}\), and combining it with Equation B.12 gives the practical test: aggregate until \(c_{X}\) falls far enough that the mass below zero is negligible, then check rather than assume.

For example, the belt’s weekly demand at \(c_{X} = 0.7449\) would put about 9% of a normal model below zero, which is not a small error in a tail but a model asserting that demand is frequently negative. Aggregated over fourteen weeks, \(c_{X} = 0.1991\) and the mass below zero is about one in four million. Notice that the family did not change and the aggregation did. The normal was wrong weekly and safe over fourteen weeks for the reason Equation B.12 gives, and the test that settles it is the coefficient of variation of the quantity actually being modeled.

Equation B.11 assumed a fixed number of terms. In an inventory system the number of terms is usually random, because demands arrive at random times, so the next two sections build the process that counts them before Section B.5 sums over it.

B.3 The Poisson Process

The sections above treated demand as a quantity. This one treats it as a stream of events arriving in time, which is what the continuous review models of Chapter 8 actually face. The Poisson process is the standard model for such a stream, and the inventory results lean on it harder than on anything else in this appendix.

Let \(N(t)\) represent the number of events in the interval \((0, t]\), with \(N(0) = 0\). The collection \(\{N(t), t \ge 0\}\) is a counting process. It is a Poisson process with rate \(\lambda\) when three properties hold.

  1. Independent increments. The numbers of events in non-overlapping intervals are independent.
  2. Stationary increments. The number of events in an interval depends only on the interval’s length and not on where it starts.
  3. Orderliness. Events do not occur simultaneously. In an interval of length \(h\) the probability of exactly one event is \(\lambda h + o(h)\) and the probability of two or more is \(o(h)\).

From those three,

\[ P\{N(t) = x\} = \frac{e^{-\lambda t}(\lambda t)^{x}}{x!}, \qquad E[N(t)] = \mathit{Var}[N(t)] = \lambda t \tag{B.15}\]

so the count over any interval is the Poisson distribution of Section C.1.1 with mean \(\lambda t\). Notice that the mean and the variance are equal, which is Equation C.1 equal to one, and that this is forced by the three properties rather than assumed separately.

You should read the three properties as assumptions to be checked and not as technical preliminaries. Independent increments is the one that fails most often in demand data, because a promotion, a shortage at a competitor, or a shared weather event correlates one interval with the next. Figure 7.3 is that check performed on a real series.

B.3.1 Interarrival Times

Let \(T_{i}\) represent the time between the \((i-1)\)st and \(i\)th events. The interarrival times of a Poisson process are independent and exponentially distributed with rate \(\lambda\), and the converse holds as well, so the two characterizations are the same process described from different ends. Thus, a mean interarrival time of \(1/\lambda\) and a mean rate of \(\lambda\) carry the same information, and which one to work with is a matter of convenience.

For example, a spare part demanded on average once every 5 weeks has \[ \lambda = \frac{1 \text{ demand}}{5 \text{ weeks}} = 0.2 \text{ demands per week} \] and over a lead time of 3 weeks the demand count is Poisson with mean \(0.2(3) = 0.6\) demands. Notice that the rate and the interval must be measured in the same time unit, which is the same caution Equation B.30 carries.

The exponential is memoryless, so the time until the next demand does not depend on how long we have already waited. That is, a Poisson process has no notion of being overdue. It is the strongest statement in this appendix and the reason the Poisson process is both tractable and often wrong.

B.3.2 Superposition, Thinning, and Conditioning

Three operations preserve the process, and all three are used later.

Superposition. Merging independent Poisson processes with rates \(\lambda_{1}, \ldots, \lambda_{n}\) gives a Poisson process with rate \(\sum\lambda_{i}\). Thus, demand at a warehouse serving \(n\) independent retailers is Poisson when each retailer’s is, which is what lets Chapter 9 treat an upstream location as facing one stream.

Thinning. If each event is independently retained with probability \(p\), the retained events form a Poisson process with rate \(p\lambda\) and the discarded events form an independent Poisson process with rate \((1-p)\lambda\). For example, if one demand in ten is for an emergency order, the emergency stream is Poisson with rate \(\lambda/10\) and is independent of the routine stream.

Conditional uniformity. Given \(N(t) = n\), the \(n\) event times are distributed as the order statistics of \(n\) independent observations drawn uniformly on \((0, t)\). That is, conditional on how many events occurred, the process says nothing about when they occurred beyond that they are scattered at random. Section B.3.3 uses this property together with thinning, and the result it gives is the most surprising one in this appendix.

B.3.3 Palm’s Theorem

Suppose events occur in a Poisson process at rate \(\lambda\), and each event starts an activity whose duration is drawn independently from a distribution with mean \(E[L]\). Let \(\mathit{IO}\) represent the number of activities in progress at a random instant in steady state. Then

\[ P\{\mathit{IO} = x\} = \frac{e^{-\lambda E[L]}\left(\lambda E[L]\right)^{x}}{x!} \tag{B.16}\]

Poisson with mean \(\lambda E[L]\), whatever the distribution of the duration is. Only its mean enters. That is Palm’s theorem, and the same result is known as the \(M/G/\infty\) occupancy distribution.

The argument uses the two properties above and nothing else. Condition on the number of events in \((0, t]\). Given that there were \(n\), their times are independent and uniform on \((0, t)\) by conditional uniformity. An event at time \(u\) is still in progress at \(t\) with probability \(1 - G_{L}(t-u)\), and those outcomes are independent across events, so the events still in progress are a thinning of the \(n\) with retention probability

\[ q = \frac{1}{t}\int_{0}^{t}\left[1 - G_{L}(v)\right]dv \tag{B.17}\]

A thinned Poisson count is Poisson with mean \(\lambda t q\), and as \(t \to \infty\) the integral in Equation B.17 tends to \(E[L]\), which gives Equation B.16. For a compound process in which each event carries a quantity, the mean becomes \(\lambda E[L] E[Y]\) by Equation B.27.

Read Equation B.16 against Equation B.30 before using either. They appear to contradict one another. Equation B.30 says that when the duration is random its variance inflates the total, through the term \(\mathit{Var}[L](E[D])^{2}\), and Equation B.16 says the variance does not matter at all. Both are correct, and they describe two different kinds of randomness.

  • Equation B.16 is about variability between activities. Each draws its own duration independently, so a long one and a short one overlap, and the long one delays nothing. Averaged over many, only the mean survives.
  • Equation B.30 is about the total accumulated over one duration. If that duration is a single unknown common to everything in progress, then a long draw stretches all of it at once and the variance is not averaged away.

Notice that the two are not alternative analyses of one situation. They are analyses of two situations, and the question that separates them is whether activities started later can finish earlier. Section 8.5.4 applies this distinction to a replenishment system, where the answer decides which formula is the right one.

B.3.4 PASTA

An inventory system is measured in two different ways and the difference matters. The ready rate is the fraction of time that stock is on the shelf. The fill rate is the fraction of demands that are filled from the shelf. The first is an average over time and the second is an average over customers, and in general they are different numbers.

They coincide under one condition. Poisson Arrivals See Time Averages, or PASTA, says that when arrivals form a Poisson process and do not anticipate the future of the system, the distribution of the system state found by an arriving customer equals the long-run time-stationary distribution. Thus, under Poisson demand the fraction of demands arriving to an empty shelf equals the fraction of time the shelf is empty, and the fill rate equals the ready rate.

Notice the size of what that assumption is doing. Under any other arrival process an arriving demand is more likely to find the system in some states than a randomly chosen instant is, and the two service measures separate. Chapter 8 keeps them apart for exactly this reason, and reports which one a formula computes.

B.3.5 Demand Over a Random Lead Time

We can now settle a question Chapter 8 will ask directly: if demand is Poisson and the lead time is itself random, what distribution does the lead time demand have?

Let \(D\) represent the demand over a lead time \(L\), where \(L\) is independent of the demand process. Conditioning on \(L = \ell\) in the manner of Equation B.7, the demand is Poisson with mean \(\lambda\ell\), so the probability generating function of \(D\) is

\[ \tilde{g}_{D}(z) = E\left[e^{-\lambda L(1-z)}\right] = \tilde{f}_{L}\big(\lambda(1-z)\big) \tag{B.18}\]

where \(\tilde{f}_{L}\) is the Laplace transform of the lead time. That is, the lead time demand is obtained by evaluating the lead time’s own transform at \(\lambda(1-z)\), and no new machinery is needed.

Take \(L\) to be Erlang with shape \(n\) and rate \(\mu\), which by Section C.2.2 is the sum of \(n\) independent exponential stages. Its transform is \(\tilde{f}_{L}(s) = \left(\mu/(\mu+s)\right)^{n}\), so

\[ \tilde{g}_{D}(z) = \left(\frac{\mu}{\mu + \lambda(1-z)}\right)^{n} = \left(\frac{p}{1 - (1-p)z}\right)^{n}, \qquad p = \frac{\mu}{\lambda + \mu} \tag{B.19}\]

which is the generating function of the negative binomial of Equation C.3 with \(r = n\) and that \(p\). Thus, Poisson demand over an Erlang lead time is exactly negative binomial. Section C.1.2 reaches for the negative binomial because real demand has \(\mathit{VMR} > 1\); Equation B.19 says where that extra variability comes from when the lead time is the source of it.

A warning about \(p\). Recall from Section C.1.2 that two conventions are in circulation. Equation B.19 is in the convention of Equation C.3, where \(p\) is the probability of success, so \(p = \mu/(\lambda+\mu)\). The failure convention gives \(\lambda/(\lambda+\mu)\) instead, and at \(\lambda = 3\) and \(\mu = 2\) those are 0.400 and 0.600. Notice that both are admissible values, so a routine handed the wrong one returns an answer rather than an error.

The arithmetic check. Take \(\lambda = 3\) demands per week, \(n = 4\) and \(\mu = 2\) per week, so \(p = 2/5 = 0.400\) and \(r = 4\). From Equation C.3, \(\beta = (1-p)/p = 1.5\), so \(E[D] = 4(1.5) = 6\) units and \(\mathit{Var}[D] = 4(1.5)/0.4 = 15\) units squared. Now compute the same two moments from Equation B.30, using \(E[L] = n/\mu = 2\) weeks and \(\mathit{Var}[L] = n/\mu^{2} = 1\) week squared:

\[ E[D] = 3(2) = 6, \qquad \mathit{Var}[D] = 3(2) + 3^{2}(1) = 6 + 9 = 15 \]

The two routes agree. That check is worth performing every time a transform argument produces a named family, because it catches an inverted convention immediately. Passing \(p = 0.600\) to Equation C.3 instead would have reported a mean of 2.667 units and a variance of 4.444, and neither figure announces itself as wrong.

The Poisson process assumes exponential interarrival times, and real demand rarely obliges. The next section keeps the counting and drops that assumption.

B.4 Renewal Processes

A renewal process keeps the counting of Section B.3 and drops the exponential interarrival time. Let \(T_{1}, T_{2}, \ldots\) represent independent and identically distributed positive interarrival times with mean \(\mu_{T} = E[T]\) and variance \(\sigma_{T}^{2}\), let \(S_{n} = T_{1} + \cdots + T_{n}\) be the time of the \(n\)th event, and let

\[ N(t) = \max\{n : S_{n} \le t\} \tag{B.20}\]

count the events in \((0, t]\). The Poisson process is the special case in which \(T\) is exponential. Thus, everything below specializes to Equation B.15, and we use that as the check on each result.

Notice the relationship Equation B.20 encodes: \(N(t) \ge n\) exactly when \(S_{n} \le t\). That is, a statement about a count is a statement about a sum, so Section B.2 supplies the distribution of \(S_{n}\) and this section translates it into statements about \(N(t)\).

B.4.1 How Many Events, and How Variable

The renewal function \(m(t) = E[N(t)]\) has no convenient closed form for most interarrival distributions, and for large \(t\) it does not need one. The elementary renewal theorem states

\[ \lim_{t \to \infty}\frac{m(t)}{t} = \frac{1}{\mu_{T}} \tag{B.21}\]

so events accumulate at the reciprocal of the mean interarrival time. For example, a part demanded on average once every 5 weeks generates about \(t/5\) demands in \(t\) weeks, which is the statement a planner would have made without any of this machinery. Equation B.21 is the guarantee that the obvious statement is correct, and that only the mean of the interarrival time enters it.

The variance is where the interarrival distribution reappears. For large \(t\),

\[ \mathit{Var}[N(t)] \approx \frac{\sigma_{T}^{2}}{\mu_{T}^{3}}\,t \tag{B.22}\]

Check it against the Poisson process, where \(T\) is exponential with \(\mu_{T} = 1/\lambda\) and \(\sigma_{T}^{2} = 1/\lambda^{2}\). Then Equation B.22 gives \((1/\lambda^{2})(\lambda^{3})t = \lambda t\), which is Equation B.15 exactly. That agreement is the check worth performing on any renewal formula, because the Poisson case is the one whose answer is already known.

Equation B.22 is a rate and not the whole answer. The exact variance differs from it by a bounded amount that does not grow with \(t\), and that correction depends on the third moment of the interarrival time and not only on its first two. That is why a routine that computes the variance of the number of demands in an interval asks for three moments of the time between demands, which will look like an error to a reader who has only seen Equation B.22. For a Poisson process the correction is exactly zero, which is another way of saying the exponential is the case in which two moments suffice.

B.4.2 A Random Horizon

Chapter 8 counts demands over a lead time, which is itself random. Let \(L\) represent the lead time, independent of the demand process. Applying Equation B.8 to \(N(L)\), with Equation B.21 for the conditional mean and Equation B.22 for the conditional variance,

\[ E[N(L)] \approx \frac{E[L]}{\mu_{T}}, \qquad \mathit{Var}[N(L)] \approx \frac{\sigma_{T}^{2}}{\mu_{T}^{3}}E[L] + \frac{\mathit{Var}[L]}{\mu_{T}^{2}} \tag{B.23}\]

Read the two variance terms the way Section B.5.2 reads its own. The first is the variability of the demand stream, accumulated over the expected length of the lead time. The second is the variability of the lead time, converted into a count by dividing by the mean interarrival time twice. Thus, a supplier who is unreliable and a demand stream that is erratic enter the same expression through different doors, and computing the two terms separately tells you which door to close.

The check. Take demand Poisson at rate \(\lambda\), so \(\mu_{T} = 1/\lambda\) and \(\sigma_{T}^{2} = 1/\lambda^{2}\). Then Equation B.23 gives

\[ \mathit{Var}[N(L)] \approx \lambda E[L] + \lambda^{2}\mathit{Var}[L] \]

which is Equation B.30 for demands of one unit each. For example, at \(\lambda = 3\) per week and an Erlang lead time with \(E[L] = 2\) weeks and \(\mathit{Var}[L] = 1\) week squared, the variance is \(3(2) + 9(1) = 15\), which is the figure Equation B.19 produced exactly. Three routes, one answer.

B.4.3 Little’s Law

One result about long-run averages holds with almost no assumptions and is used throughout Chapter 8. Let \(\lambda\) represent the long-run arrival rate into a system, let \(\overline{W}\) represent the average time an arrival spends in it, and let \(\overline{L}\) represent the average number in it. Then

\[ \overline{L} = \lambda\overline{W} \tag{B.24}\]

Little’s Law requires only that the averages exist and that the system is in steady state. It does not require Poisson arrivals, independence, or any distributional assumption, and that generality is what makes it useful.

Its inventory reading is Equation 1.6, which Section 1.6 obtained by this argument before the machinery for it existed. Take the system to be the set of backordered demands. Then \(\overline{L}\) is the average backorder level \(\overline{B}\), \(\lambda\) is the demand rate, and \(\overline{W}\) is the average time a demanded unit waits. Thus, a constraint on the average number of units backordered is the same constraint as one on how long a customer waits, and two service targets an organization would argue about separately are one target in two units. Chapter 8 uses Equation 1.6 in both directions.

B.4.4 Demands That Arrive in Batches

Nothing so far has said how much is demanded at an event. Let \(D_{i}\) represent the size of the \(i\)th demand. The total demand over an interval of length \(t\) is

\[ \sum_{i=1}^{N(t)} D_{i} \tag{B.25}\]

a compound renewal process, and a random sum in which the count comes from Equation B.20. When the counting process is Poisson this is the compound Poisson of Section C.4.3.

Equation B.25 is the object Section B.5 takes up, and it now has both of its ingredients: this section supplies \(E[N(t)]\) and \(\mathit{Var}[N(t)]\), and the next supplies what to do with them once the sizes are random too.

B.5 Random Sums

A random sum adds up a random number of random quantities. Let \(N\) represent a non-negative integer random variable, let \(X_{1}, X_{2}, \ldots\) represent identically distributed random variables with the same distribution as some \(X\), and let

\[ S = \sum_{i=1}^{N} X_{i}, \qquad S = 0 \text{ when } N = 0 \tag{B.26}\]

Assume throughout that the \(X_{i}\) are independent of one another and of \(N\). That last requirement is the one that fails in practice, and Section B.5.5 says what it costs.

Random sums appear twice in this book, and both times the quantity being summed is demand.

Demand over a random lead time. Chapter 8 needs the demand that occurs while a replenishment order is outstanding. Let \(D_{i}\) represent the demand in period \(i\) and let \(L\) represent the lead time measured in whole periods. Then the lead time demand is \(D(L) = \sum_{i=1}^{L}D_{i}\), a random sum in which \(N\) is the lead time.

Demand that arrives in batches. A slow-moving item is not demanded every period. Let \(N(t)\) represent the number of demand occurrences in an interval of length \(t\) and let \(D_{i}\) represent the size of the \(i\)th one. Then the demand over the interval is \(\sum_{i=1}^{N(t)}D_{i}\), a random sum in which \(N\) counts occurrences rather than periods. Section C.4.3 uses this representation for intermittent demand.

B.5.1 The Mean

Condition on \(N\). Given \(N = n\) the sum has exactly \(n\) terms, so \(E[S \mid N = n] = nE[X]\). Thus, \(E[S \mid N] = NE[X]\), and taking the expectation of both sides,

\[ E[S] = E\big[E[S \mid N]\big] = E\big[NE[X]\big] = E[N]\,E[X] \tag{B.27}\]

Equation B.27 is Wald’s equation, and it says the obvious thing: the expected total is the expected count times the expected size. Notice that it needs only the means of \(N\) and \(X\), which is why it is the one result in this section that survives almost any weakening of the assumptions.

B.5.2 The Variance

The variance is where the interesting behavior is, and the obvious guess is wrong. Conditioning again, and using the law of total variance,

\[ \mathit{Var}[S] = E\big[\mathit{Var}[S \mid N]\big] + \mathit{Var}\big[E[S \mid N]\big] \tag{B.28}\]

Take the two terms in turn. Given \(N = n\) the sum is of \(n\) independent terms, so \(\mathit{Var}[S \mid N] = N\mathit{Var}[X]\) and the first term is \(E[N]\mathit{Var}[X]\). And \(E[S \mid N] = NE[X]\), so the second term is \(\mathit{Var}[NE[X]] = (E[X])^{2}\mathit{Var}[N]\). Thus,

\[ \mathit{Var}[S] = E[N]\,\mathit{Var}[X] + \mathit{Var}[N]\,\big(E[X]\big)^{2} \tag{B.29}\]

The two terms are not symmetric, and that asymmetry is the point. The variability of the sizes enters multiplied by the expected count. The variability of the count enters multiplied by the square of the expected size. That is, uncertainty in how many contributes a term that grows quadratically in the typical size, while uncertainty in how much contributes only linearly in the typical count.

For example, an item demanded five units at a time is four times as exposed to uncertainty in the number of demands as an item demanded two and a half units at a time, all else equal. Recall from Section 8.3.2 that a supplier’s reliability and a customer’s predictability are different problems. Equation B.29 is why they are also different sizes of problem.

B.5.3 Lead Time Demand

Specializing Equation B.27 and Equation B.29 to \(D(L) = \sum_{i=1}^{L}D_{i}\) gives the two results Section C.4.2 and Chapter 8 both use:

\[ E[D(L)] = E[L]\,E[D], \qquad \mathit{Var}[D(L)] = E[L]\,\mathit{Var}[D] + \mathit{Var}[L]\,\big(E[D]\big)^{2} \tag{B.30}\]

Notice the units. \(E[L]\) is in periods and \(E[D]\) in units per period, so the product is in units. In the variance, \(E[L]\mathit{Var}[D]\) is periods times units squared per period squared, and \(\mathit{Var}[L](E[D])^{2}\) is periods squared times units squared per period squared. Both come to units squared, which is the check that the two terms belong in the same sum. Notice also that \(L\) and \(D\) must be measured in the same period. A lead time in days and a demand per week is the most common error in applying Equation B.30, and it is invisible in the answer.

For example, take an item averaging 5 units a week with a variance of 16, and a supplier whose lead time averages 10 weeks with a variance of 25. Then

\[ E[D(L)] = 10(5) = 50 \text{ units} \]

\[ \mathit{Var}[D(L)] = 10(16) + 25(5)^{2} = 160 + 625 = 785 \text{ units}^{2} \]

so the standard deviation of lead time demand is \(\sqrt{785} = 28.02\) units.

Read the two terms before moving on. The demand contributes 160 of the 785 and the lead time contributes 625, so about 80% of the variance the safety stock must cover comes from the supplier rather than from the customer. That split is typical, and it is the argument for spending on lead time reliability before spending on forecasting. You should compute both terms separately every time, because the total alone does not say where to act.

B.5.4 Batches, and the Compound Poisson

When \(N\) counts demand occurrences rather than periods, Equation B.29 applies unchanged with \(N = N(t)\). What changes is that \(E[N(t)]\) and \(\mathit{Var}[N(t)]\) now have to come from somewhere, and Section B.4 is where.

One case is clean enough to state here. If demands occur according to a Poisson process with rate \(\lambda\), then \(N(t)\) is Poisson with \(E[N(t)] = \mathit{Var}[N(t)] = \lambda t\), and the two terms of Equation B.29 collapse:

\[ \mathit{Var}[S] = \lambda t\left(\mathit{Var}[X] + (E[X])^{2}\right) = \lambda t\,E\big[X^{2}\big] \tag{B.31}\]

Thus, a compound Poisson demand needs only the second moment of the batch size, not its variance and mean separately. For example, at three demand occurrences on average, each averaging 5 units with a variance of 16, the total averages \(3(5) = 15\) units with a variance of \(3(16) + 3(25) = 123\), which Equation B.31 gives directly as \(3(41) = 123\). The two routes agree, and that agreement is the arithmetic check for this section.

Notice what Equation B.31 implies. The variance to mean ratio of a compound Poisson demand is \(E[X^{2}]/E[X]\), which exceeds one whenever the batch size is ever larger than one unit. That is the reason Section C.1 sends batched demand to the negative binomial rather than the Poisson, and it is why a Poisson process does not produce Poisson period demand unless every demand is for a single unit.

B.5.5 What This Does and Does Not Give You

Equation B.27 and Equation B.29 give two moments. They do not give a distribution, and the inventory models of Chapter 8 onward need one. Section C.4 is the bridge, and it is an assumption rather than a derivation: two numbers do not determine a distribution, so a family has to be chosen.

Two assumptions stand behind the results above, and both fail in real data.

The count must not depend on the sizes. Suppose a large demand triggers an expedite that shortens the lead time. Then \(L\) and the \(D_{i}\) are dependent. Equation B.27 survives that under weaker conditions; Equation B.29 does not, and the true variance is usually smaller than it gives.

The sizes must not depend on one another. Demand that is autocorrelated, a run of heavy weeks followed by a run of light ones, breaks the first term of Equation B.29. Recall from Section 2.2.6 that burstiness is one of the things a demand history is characterized for. An item that tests as bursty is an item for which Equation B.29 understates the spread, and understating the spread understates the safety stock.