Record Accuracy: The Number the Policy Actually Reads

Manuel D. Rossetti, PhD P.E.

Agenda

  • Why records drift, and what that does to a correct policy
  • What error actually costs, and why nobody notices
  • Three definitions of accuracy, one count sheet, three answers
  • Cycle counting: methods, frequency, and what counting is really for
  • Excess and obsolete stock

Nobody Sets a Reorder Point by Walking to the Shelf

The policy reads a record.

A record is a claim about the shelf, not the shelf itself.

Claims decay.

Every policy in this course triggers on a number the system believes.

What a Record Claims

At a minimum, four things:

  • a stock number
  • a location
  • a quantity on hand
  • a condition

The record is inaccurate if any one of them is wrong.

That is stricter than it first appears, and it is the right definition. An item recorded in the correct quantity at the wrong location cannot be picked. An item in the right place with the right count but the wrong condition code will be issued to a customer who cannot use it.

Where Errors Come From

Source What happens Its signature
Transaction error A receipt or transfer keyed with the wrong quantity A discrepancy with no offsetting entry
Unrecorded movement Stock taken or returned without a transaction Record above shelf, drifting one way
Shrinkage Theft, breakage, spoilage Record above shelf
Misidentification Two similar items confused Paired errors, one high and another low
Unit of issue Each against case, pack against piece Errors that are exact multiples
Timing Count and transaction on opposite sides of a cutoff A discrepancy that resolves itself
Location error Stock moved without a transfer Correct total, wrong place

The signature matters because it is what makes a cause findable.

Two Mechanisms, Two Responses

Transaction errors occur when a movement is recorded incorrectly. They are process defects and can be designed out.

Stock loss errors are unrecorded losses through shrinkage, damage, or theft, where the movement was real and no transaction was ever generated. No amount of transaction discipline prevents them, and they drift in one direction, with the record always above the shelf.

Notice one feature that organizes the whole subject: every source is a transaction, or the absence of one.

A record for an item that nothing has happened to cannot become wrong.

The Quantity Is Actually Two Quantities

  • \(I_{a}(t)\), the amount actually on the shelf
  • \(I_{r}(t)\), the amount the system believes is there

\[ D(t) = I_{a}(t) - I_{r}(t) \]

Shrinkage and unrecorded issues make \(D(t)\) negative: the record stands above the shelf.

The two are used for different things. Demand is filled from \(I_{a}(t)\), because that is what physically exists. But the reorder decision is made on

\[ \mathit{IP}(t) = I_{r}(t) + \mathit{IO}(t) - B(t) \]

How It Fails, and Why Nobody Sees It

A continuous review system triggers when the recorded position reaches \(r\).

If the record overstates what is on the shelf, the trigger fires late.

The effective reorder point is lower than the one designed, and the item runs short more often than the design allows.

Nothing in the system reports this.

The model reports the service it was designed to deliver. The warehouse delivers something worse.

And You Pay for Safety Stock Twice

Safety stock exists to absorb uncertainty in demand over the lead time.

Record error is a second source of uncertainty, and the safety stock absorbs it too, whether or not anyone intended that.

An organization with poor records is paying for safety stock twice: once for the demand uncertainty it knows about, and again for the record uncertainty it does not.

What Error Costs: Three Responses

Simulation of a two-echelon chain, each location running an \((r,Q)\) policy, separates what an organization can do about wrong records.

Leave the policy alone. In the high error case, roughly half the fill rate was lost, and system backorders rose from about ten items to about three hundred and seventy.

Re-optimize around the error. Recomputing \(r\) and \(Q\) to hit the same 90% fill rate does restore the service, and pays for it in stock. Backorders still rise sharply, by roughly 300% for fast movers.

Count. Adding cycle counting raised fill rates while holding inventory the same or slightly lower.

The Second Response Explains Everything

Re-optimizing around the error is why record inaccuracy survives for years without being raised.

The customer-facing measure is back at target. The cost has moved into inventory, transportation, and handling, where nobody attributes it to record keeping.

The symptom that would have prompted an investigation has been bought off.

Counting is not a trade. The first two responses trade service against inventory; counting improves service without buying it with stock, because it attacks the discrepancy instead of compensating for it.

Measuring Accuracy: A Deceptively Simple Ratio

\[ \text{record accuracy} = \frac{\text{number of accurate records}}{\text{number examined}} \]

That hides three choices, and each changes the answer.

What counts as accurate. Exact match, or a tolerance? Exact agreement on a bin of ten thousand washers is not worth what it costs to obtain.

What is weighted. Every record equally, or weighted by value?

Which fields. Quantity alone is the usual reported figure. Location and condition accuracy are separate measures.

One Count Sheet

Item Unit cost Recorded Counted Error
A-100 $80.00 120 120 0.0%
A-101 $150.00 45 44 2.2%
A-102 $2.00 500 495 1.0%
B-200 $400.00 12 12 0.0%
B-201 $0.40 3,000 2,970 1.0%
B-202 $25.00 8 0 100.0%
C-300 $6.00 250 250 0.0%
C-301 $0.75 1,000 990 1.0%
C-302 $3.00 60 66 10.0%
C-303 $95.00 15 15 0.0%

Three Answers

Strict match. Four records agree exactly:

\[ \frac{4}{10} = 40\% \]

Within a 2% tolerance. The four exact ones, plus A-102, B-201, C-301 at 1.0%:

\[ \frac{7}{10} = 70\% \]

Weighted by value. Recorded value $27,405, absolute discrepancy $397.50:

\[ 1 - \frac{397.50}{27{,}405} = 98.55\% \]

The Same Shelf: 40%, 70%, or 98.6%

The value-weighted figure is highest because the errors sit on cheap items.

Now look at what the dollar figure does to B-202. The item is missing entirely, eight units recorded and none on the shelf. It moves the value-weighted accuracy by $200 out of $27,405.

The next customer who wants B-202 will not care that the warehouse is 98.6% accurate by value.

Aggregate dollar value masks individual errors, which is the central criticism of accuracy programs built on financial sampling. Picking an item requires that item’s record to be right.

Cycle Counting, and What It Is For

An annual physical inventory counts everything at once. As a way of managing accuracy it has three defects: it is disruptive, it is infrequent, and it diagnoses nothing.

Cycle counting counts a portion continuously, on a schedule, without stopping operations. Its purpose, in order:

  1. Identify the causes of errors
  2. Correct the conditions causing them
  3. Maintain a high level of record accuracy
  4. Provide a correct statement of assets

Accuracy is third on that list and the balance sheet is fourth.

A program judged only on records corrected is being judged on its by-product.

Methods, and Two That Look Wrong

Method How the sample is formed
Random sample Drawn at random; statistically clean, ignores everything you know
ABC Stratified by class, each counted at its own frequency
Process control The counter counts what is easy to count and skips the rest
Location based An area is chosen, everything in it counted
Opportunity based Triggered by events: a reorder, a put-away, a balance near \(r\)
Transaction based Counted after every \(n\) transactions

Process control counting is deliberately biased, and its advocates know it. The argument: the bias runs toward the items most likely to be wrong. If the purpose is to find causes rather than estimate a population parameter, a biased sample aimed at the problem beats an unbiased one that is mostly clean records.

Transaction based counting follows from the fact that errors enter through transactions. Counting an item that has not moved since it was last counted correct is close to guaranteed to find nothing.

Opportunity Counting: Three Free Triggers

Ordinary operation exposes discrepancies for nothing:

  • A demand arrives, the record says there is not enough, the stock is visibly on the shelf. The record is understating.
  • A demand arrives, the record shows a positive balance, the shelf cannot fill it. The record is overstating, and the stockout just made that visible.
  • The recorded balance approaches the reorder point.

The third is the most useful, because it verifies the record precisely when a wrong one would cause a wrong order.

How Often Should You Count?

Let \(p\) be the probability a verified record goes wrong in one period, and \(n\) the periods between counts. Average accuracy over the cycle is

\[ A(n) = \frac{1}{n}\sum_{t=0}^{n-1}(1-p)^{t} = \frac{1-(1-p)^{n}}{n\,p} \]

With \(p = 0.03\) and a 95% target:

\(n\) months 1 2 3 4 5 6
\(A(n)\) 100.0% 98.5% 97.0% 95.6% 94.2% 92.8%

Count every four months, three times a year.

Double the Error Rate

With \(p = 0.06\): \(A(2) = 97.0\%\) and \(A(3) = 94.1\%\).

So it needs counting every two months, six times a year, to hold the same target.

Notice what did not enter that calculation: annual dollar volume.

It is \(p\), not value, that determines how often a record needs verifying. Where the two disagree, counting by value alone spends effort in the wrong place.

And they do disagree. A field experiment found low value, high volume items showing the lowest accuracy. That is precisely the population an ABC scheme counts least often.

But Accuracy Is Not the Same as the Value of Accuracy

Simulation of a retail chain found cycle counting

  • essential for low demand, high cost items, producing substantial savings
  • close to worthless for high demand, low cost items

Set that beside the field experiment: the cheap fast items have the worst records and are the least worth correcting, because the consequence of an error on them is small and a little extra stock covers it.

Counting effort should follow the cost of being wrong, not the probability of being wrong.

Sizing a Sample

Estimating population accuracy is a separate problem from maintaining it.

\[ m = \frac{z_{\alpha/2}^{2}\,\hat{a}(1-\hat{a})}{e^{2}} \]

Accuracy near 92%, within two percentage points, 95% confidence:

\[ m = \frac{(1.96)^{2}(0.92)(0.08)}{(0.02)^{2}} = 706.9 \]

707 records. Tighten to one percentage point and it quadruples to 2,828, because \(m\) varies as \(1/e^{2}\).

That is why accuracy is usually reported to the nearest percentage point and no finer.

Monitoring Over Time

Accuracy is an attribute: each record sampled is accurate or it is not. A sample of \(n\) yields a proportion, which is exactly what a \(p\) chart consumes.

With \(n = 250\) and \(\bar{p} = 0.92\):

\[ s = \sqrt{\frac{(0.92)(0.08)}{250}} = 0.0172 \]

\[ \text{limits} = 0.92 \pm 3(0.0172) \quad\Longrightarrow\quad [0.869,\ 0.972] \]

Reading It

This week: 214 of 250 accurate, or 85.6%. Below the lower limit. The process has changed; investigate.

Contrast a week returning 220 of 250, or 88.0%. Four points below the center line, and it looks alarming.

But it sits inside the limits. A sample of 250 drawn from a process running at 92% produces readings that low often enough that one of them means nothing.

Acting on it would send a team looking for a cause that does not exist.

Two Cautions on the Chart

The population is not homogeneous. A \(p\) chart assumes every unit sampled comes from the same process with the same probability of conforming. Inventory records do not satisfy that: each item has its own propensity to go wrong, so the false alarm rate is inflated. Simulation found the inflation slight, but the assumption is violated.

A signal is a prompt to investigate, not to adjust. An out-of-control point says the process producing records has changed. Correcting the sampled records addresses none of that.

The chart’s real value is not detecting deterioration. It is demonstrating stability, and showing whether an improvement actually moved the level.

Correction Is Not the Point

A count that finds a discrepancy and adjusts the record has closed a symptom.

If the condition that produced it is still in place, the same record will be wrong again, and the program will count the same errors indefinitely at a steady cost.

The most useful finding in the practitioner literature is a negative one. Benchmarking more than twenty companies, best-in-class accuracy of 99% and above was achieved by organizations that dedicated resources, had strong system support, and concentrated on eliminating common process errors.

They did not emphasize statistical methods.

Read That Carefully

You should know the statistical material. It tells you how often to count, what a measured accuracy figure does and does not support, and how large a sample an estimate requires.

What the benchmarking says is that statistical sophistication is not sufficient and is not the binding constraint.

An organization counting diligently on crude rules and fixing what it finds will outperform one sampling optimally and adjusting records without asking why they were wrong.

Accuracy Is Not a Local Property

A location’s replenishment orders are the demand its supplier sees.

A retailer ordering against bad records sends distorted demand upstream.

In the two-echelon experiments, a single retailer declining to cycle count, while every other location did, cost that chain roughly 6% of fill rate, added about 2% to inventory, and raised backorders by roughly 400%.

The cost of one partner’s records lands on partners who have no control over them.

Which is why record accuracy belongs in the performance measures a supply chain agrees on, and, between separate firms, in the contract.

Excess and Obsolete Stock

Excess is more than will be used in any reasonable horizon. Obsolete will not be used at all.

It accumulates because every model in this course assumes demand continues. A replenishment policy has no mechanism for noticing that demand has stopped. It sees position fall below \(r\) and orders, exactly as designed, until someone intervenes.

Excess is not a failure of the policy. It is the absence of a decision the policy was never asked to make.

Keeping Stock That Will Never Be Used

400 units at $28 each. Annual demand has fallen to 12 units, so the stock is more than 33 years of supply. \(i = 0.20\), \(f = \$30\) per year. A liquidator offers $6 per unit.

Retain two years, 24 units, leaving 376 in excess. With \(h = (0.20)(28) = \$5.60\):

\[ (376)(5.60) + 30 = \$2{,}135.60 \text{ per year to keep it} \]

\[ (376)(6.00) = \$2{,}256 \text{ now, to dispose of it} \]

One year of holding costs very nearly the entire salvage value. The break-even is about 1.06 years, against three decades of demand.

Why This Decision Is Hard

The number that makes it hard does not appear above.

The excess cost \((376)(28) = \$10{,}528\) to buy. Disposing at $6 realizes a write-off of over $8,000 that someone must sign.

That $10,528 is sunk. It is gone whether the stock is sold, held, or scrapped, and it is not a reason to keep anything.

But the write-off is visible and the $2,135 a year is not.

So the common outcome is to spend the larger amount quietly rather than recognize the smaller one publicly.

What We Did Today

A policy reads a record, not a shelf, and the reorder decision uses the record while demand is filled from the shelf.

Counting is not a trade against inventory. Giving up service or holding more stock both mask the problem; counting removes it.

Accuracy is not one number. One count sheet gave 40%, 70%, and 98.6%.

Count frequency follows the error rate, and counting effort should follow the cost of being wrong.

Correction is the by-product. A program that only adjusts will find the same errors forever.

For Next Time

Read before next session: The General Cycle, and Average Inventory and Backorders.

We are done with the prerequisites. Everything from here answers how much and when, for one item, beginning with the strongest assumption available: demand known and constant.

Next session:

  • One cycle, four segments, and two rates
  • Average inventory and average backorders, as triangle areas
  • Four classical models, each one assumption away from a single general result

Bring graph paper, or the ability to imagine it. Next session is geometry before it is algebra.

⌂ Index