Tool Change Intervals, Audited Against Measured Wear
A 16-insert audit of a fixed change interval against wear that was actually measured: the rule misses by about a quarter of the change point, the range a shop would quote around it covers 68% rather than the promised 90%, and pulling early enough to be safe is paid for in tool life.
Version history
- v0.1 · 2026-08-23 · working draft for review
- v0.3 · 2026-08-23 · restructured after a cold persona read: problem stated before method, the split established where it is introduced, the instrumentation findings moved forward out of the limits section, the leakage argument compressed and delegated to background/clustered-data, and the standard-practice claim grounded
- v0.3 · 2026-08-23 · adversarial review: separated the in-sample Gaussian from a held-out one, switched every interval to a cluster bootstrap over inserts, corrected the Mondrian mechanism, re-indexed the decision simulation onto physical cuts, repaired the fixed-N policy
- v1.0 · 2026-09-02 · first public release. Content unchanged from 0.3; the version marks the paper as issued rather than circulating for review.
Most machining shops change inserts on a fixed cut count. That count is a model, and unlike every other model in the building it has never been asked how often it is right. We audit it on the NASA Milling Data Set: 16 inserts run and measured, flank wear read by microscope between cuts. Three results. First, the fixed rule carries a leave-one-insert-out error of 0.165 mm, about a quarter of a conventional 0.60 mm change point. Second, the interval a shop would quote covers 68.5% against a promised 90%, and the split of blame matters more than the number: roughly 16 points come from measuring error on data the model was fitted on, only about 6 from assuming a normal shape. Third, holding out rows instead of whole inserts reports 89.8% against a nominal 90%, passes any marginal check, and is 19% too narrow while under-covering exactly where the decision is made.
- Calibration background/calibration · v1.1 · tutorial
- Clustered Data background/clustered-data · v0.1 · tutorial
Real measurements rarely arrive one at a time. They arrive twenty to a batch, nine to a tool, forty to an operator's shift. Three things change when they do. The sample is smaller than the row count says, by a factor of 1 + (m - 1) times the intraclass correlation. Any interval computed as though the rows were independent is too narrow, and the fix is to resample groups rather than rows. And any validation split that divides rows rather than groups puts the same group on both sides of the split, which produces a model evaluation that looks correct and is not. The last of these is the dangerous one, because it does not fail loudly: it reports near-nominal performance and narrower intervals, and passes the check almost everyone runs.
A number in your shop decides when inserts get pulled. 20 parts, or 40, or whatever the last argument about scrap settled on. It came from experience. Nobody has ever been able to say how right it is, because checking means stopping the machine and putting the tool under a microscope.
In 1996 two researchers at Berkeley did exactly that, 167 times. They ran 16 inserts on a milling centre across cast iron and stainless, stopped at intervals, measured the flank wear on each one, and published the lot. That is an unglamorous dataset and an unusual one: it is a tool-change decision with the right answer written down next to it.
So this article asks the question a shop floor cannot. Predict wear the way a shop implicitly does, from cut count and cutting conditions: how wrong are you, and does the range you’d quote around that prediction mean what it says?
The findings run in the order they matter to the decision: how wrong the current rule is (§ 02), an instrument that fails exactly where the decision lives (§ 03), what the range around a prediction actually covers (§ 04), and what pulling early enough to be safe costs in tool life (§ 05). § 06 is what this cannot tell you, which on a dataset of 16 tools is a long list.
The Decision, and a Dataset That Knows the Answer
Sixteen inserts, run and measured rather than run and guessed at. That is what makes the audit possible.
Tool wear data usually comes from a process log, where the wear is inferred after the fact from when somebody changed the tool. That records a judgement, not a measurement; you can’t audit a prediction against somebody’s past judgement.
The NASA Milling Data Set is an experiment instead. A Matsuura MC-510V machining centre, a 70 mm face mill carrying 6 coated carbide inserts, a constant 200 m/min. Every combination of 2 materials (cast iron and J45 stainless), 2 depths of cut (0.75 and 1.5 mm) and 2 feeds (0.25 and 0.5 mm/rev), each run twice with a fresh set of inserts. 16 inserts, 167 cuts, and after most of those cuts the tool came off the machine and its flank wear went under a microscope. 146 cuts carry a measurement. 10 of the 16 inserts reached 0.60 mm of wear; the other 6 stopped short of it.
So for every cut there are 2 things: what a prediction would have said, and what the wear actually was.
Flank wear (VB) is the width of the worn band on the cutting edge, in millimetres. It grows with use, roughly and unevenly, and past some value the surface finish and the dimensional tolerance start to go. A change point is the value at which the insert gets pulled. 0.60 mm is a conventional figure for this class of carbide, and § 06 reports every result across change points from 0.30 to 1.00 mm rather than resting on that one.
The unit here is the insert, not the cut. That one fact governs every method below. 9 to 20 successive cuts share one edge and one accumulated wear history, so 146 cuts are nowhere near 146 independent observations. Every model below is fitted on some inserts and tested on inserts it has never seen, and every range is calibrated on a third set of inserts held out from both. Why that matters, and what goes wrong when people split the rows instead, is background/clustered-datav0.1.
What the Current Rule Actually Knows
Fit the shop's rule as a model and it misses by 0.165 mm, on a decision made at 0.60.
The rule a shop runs is a model, even though nobody calls it one: it takes the cutting conditions and the count of parts cut, and returns a decision about a tool nobody has looked at. So it can be fitted, and it can be scored.
Fitted here means predicting flank wear from the cut number and the cutting conditions and nothing else, because a shop running a fixed interval has nothing else. Scored means tested on an insert the model has never seen. On that basis it misses by 0.165 mm on average.
Hold that next to the decision it feeds. The change point is 0.60 mm, so a typical error is about a quarter of the limit, on a quantity you can’t check without stopping the machine.
One insert, cut by cut. Dots are flank wear measured under a microscope with the tool off the machine; the line is what the cut-count model predicts for an insert it has never seen; the shaded band is that model's claimed 90% interval. The dashed rule is a 0.60 mm change point. Switch the construction and watch the band, not the line: the prediction never moves.
2 things in that figure are worth a click before reading on.
First, switch between the 2 inserts run under identical conditions. Same material, same depth, same feed, same machine, a fresh set of edges. They do not wear the same way. That spread between 2 nominally identical tools is not error in the model; it is the spread any honest range has to be wide enough to contain.
Second, switch the construction and watch the shaded band while ignoring the line. The prediction never moves. Every disagreement between methods is a disagreement about how much you are willing to admit you don’t know.
The people who built the dataset said as much in 1996, and it reads as a concession rather than a sales line:
Slight variations in a seemingly same setting vary the tool life considerably. All of these influences may be small, but they can add up over the period of machining to cause considerable uncertainty in predicting a priori the tool life.
A fixed change interval is not a rule of thumb standing in for a model; it is a model whose error nobody has measured.
The Sensor That Goes Blind When the Tool Is Worn
Spindle current is the obvious thing to watch. In this data it saturates, and it saturates in the high-wear runs rather than at random.
The rig recorded 6 channels: spindle motor current in both AC and DC form, vibration at the table and the spindle, and acoustic emission at both. The DC spindle current is the one an engineer would reach for first, because current tracks torque and torque tracks how hard the tool is working.
It clips at the converter’s 10 V ceiling in 44 of the 167 runs, and 35 of those are more than half clipped. It does not clip at random:
| median flank wear | |
|---|---|
| runs where DC current is more than half clipped | 0.500 mm |
| runs where it is not | 0.230 mm |
A worn insert needs more torque and draws more current, until the channel hits its ceiling. Here is the awkward part: the sensor that most directly measures cutting load goes blind exactly where the change decision is hardest. It is a fuel gauge that stops reading in the last quarter of the tank. Those cuts are ticked in madder along the bottom of Fig. 01.
This is not a defect peculiar to a 1990s rig. Any channel with a fixed range saturates at one end, and instruments get scaled to normal operation, so the end it saturates at is the end you wanted. A monitoring system that looks well behaved across a year of production is blind in precisely the minority of hours anyone wants to ask it about.
It is also invisible in a summary statistic. Averaged across a run, a clipped channel looks like a slightly high, slightly quiet one.
A second fault turned up in the same pass. 2 runs carry converter corruption, both the first run of their case, which suggests something in the acquisition startup. One of them holds values above 1e29 on channels amplified into a plus or minus 5 V range, across a fifth of its record. Those samples are masked rather than averaged through, and the run is kept and flagged rather than quietly dropped. Every number in this article was recomputed with it excluded, and the largest change anywhere is 0.30 of a percentage point.
Neither fault took more than a morning to find, and neither is visible in any summary of the data. A summary of the data is not a check on it.
Two Ways to Be Wrong About Your Own Error
A range built from the model's own errors covers 68% of the time against a promised 90%. Most of that gap is not the bell curve everyone argues about.
A prediction of 0.42 mm is not a decision input. What a decision needs is a range: 0.42 give or take how much. There is a standard way to produce one, and it is the same everywhere from spreadsheets to production software. Take the errors the model made, compute their standard deviation, and quote the prediction plus or minus 1.645 of those for a 90% range. 1The 1.645 is where the normal distribution puts 90% of its mass in a symmetric interval. That is the only place the bell curve enters, and § 04 shows it is not the main problem.
The question nobody asks out loud is which errors, and there are 2 answers that sound equally sensible.
The first is to use the errors the model made on the data it was built from. This needs no extra data and is what happens by default in most tools. It covers 68.5% of outcomes on inserts the model has never seen, against a promised 90%.
The second is to set some inserts aside, never fit on them, and use the errors there. That covers 84.4%.
The difference between those is not a subtlety about ranges. A model has partly memorised the data it was built from, so its errors there are smaller than its errors on anything new, and a range scaled to them is too narrow before any assumption about distributions has been made.
Achieved coverage against a promised 90%, measured on inserts the model never saw. The dot is the estimate; the rule through it is a 95% interval from a cluster bootstrap over inserts, which is wide because sixteen inserts is a small number of inserts. The madder segment is the gap between what was promised and what was delivered. Where an interval crosses the dashed rule, this data cannot tell you the construction misses its promise.
| How the range was built | All cuts | 95% interval | Worn tools | Median width |
|---|---|---|---|---|
| From errors on its own training data | 68.5% | [54.7, 82.4] | 30.0% | 0.260 mm |
| From errors on held-out inserts | 84.4% | [74.4, 93.8] | 55.0% | 0.465 mm |
| Same held-out inserts, no bell curve assumed | 92.2% | [86.4, 97.5] | 68.0% | 0.661 mm |
Of the 21.5 points missing from the first row, about 16 come from measuring your error on data you already used, and about 6 from assuming a bell curve. The expensive habit is not the distribution nobody can defend; it is checking your homework against the answers you copied it from.
The intervals in that table are wider than most write-ups would show, because 146 cuts came from 16 inserts and the arithmetic has to account for that. Read honestly, they say 2 different things. The first row fails unambiguously: even its optimistic end, 82.4%, falls short of the promise. The second row’s overall shortfall is not established, since its interval reaches 93.8% and includes 90.
Here is where the second row fails: on worn tools, which is the only place anyone consults it. There it covers 55.0% while the distribution-free construction covers 68.0%. A range can look adequate across the whole job and still be wrong in the last hour of the tool’s life.
The first row’s range is 2.5 times narrower than the third’s and looks far more useful on a screen. It is narrower because it is wrong.
There is a meaner version of this. Feed the model more sensor channels and its range gets narrower still, because a model with more inputs memorises harder: coverage falls from 68.5% on 4 inputs to 54.1% on 22. The held-out constructions barely move over the same range. Adding data can improve how a range looks and degrade what it means, at the same time, with no warning.
One choice decides every number above. All the held-out constructions here set aside whole inserts. Setting aside random rows instead leaves cuts from the same insert on both sides of the split, and the result is a range that reports 89.8% against a promised 90%. It passes any check anyone runs. It is 19% narrower than it should be, and it under-covers on exactly the tools you built it for. It fails by looking confident. The mechanism, and how to spot it in your own data, is background/clustered-data§ 03v0.1.
What the Trade-off Costs, in Cuts
Nearly eliminating over-limit cuts costs about 42 cuts of tool life in every 100. That is the price of certainty, in the units the stores counts.
Percentages do not settle a change interval. 2 counts do, and a shop already tracks both: cuts of tool life thrown away by pulling early, and cuts taken after the wear limit was passed.
Before the numbers, one property of the model that shapes them. It cannot predict the top of the wear range. Measured flank wear here reaches 1.53 mm; the model’s predictions never exceed 0.73 mm and never reach 0.75 mm at all, because a model of this kind averages over the examples it was trained on and cannot reach past them. Any policy threshold set above 0.73 mm therefore never fires, which is why the curve below goes flat at the right-hand end. That is a fact about the model rather than about the tools.
Move the threshold. Life thrown away counts cuts an insert still had in it when the policy pulled it; run past the limit counts cuts taken after the wear limit was crossed. Both are per 100 cuts, simulated insert by insert on inserts the model never saw. The rate boxes are yours to fill: we supply the counts, you supply what they are worth.
Both boxes start at placeholder values. They are not our estimates of anything; put your own numbers in and the ranking of policies below will change with them.
| policy | life thrown away | run past limit | cost, your units |
|---|---|---|---|
| act on the interval bottom | 0.0 | 6.0 | 1200 |
| act on the normalised bottom | 0.0 | 6.0 | 1200 |
| act on the interval top | 42.0 | 0.4 | 920 |
| act on the normalised top | 40.6 | 0.4 | 892 |
| threshold at 0.60 mm | 6.0 | 3.0 | 720 |
At a 0.60 mm change point, per 100 cuts:
| Policy | Life thrown away | Run past the limit |
|---|---|---|
| Pull when the top of the range reaches the limit | 42.0 | 0.4 |
| Pull when the prediction itself reaches its best threshold | 1.2 | 3.0 |
| Pull after a fixed 9 cuts, the best fixed count here | 22.9 | 1.8 |
| Pull only when the bottom of the range reaches the limit | 0.0 | 6.0 |
Reading the first and last rows together gives the price of near-certainty: driving over-limit cuts to nearly nothing costs about 42 cuts of tool life in every 100.
The fixed-count row needs a caveat rather than a headline. The median insert here lasted 8 cuts, so any count much above 9 simply never fires; counts of 10 and above leave the rule inert on more than half the inserts and are excluded from the comparison rather than allowed to win it by doing nothing.
No currency appears in that figure or in this section, and the omission is the point. The counts are measurable and ours to supply. What an insert costs you, what an hour of that spindle is worth, and what a part cut with a worn tool costs when it reaches your customer are all things you know and an outsider does not. A return-on-investment figure built on an outsider’s guesses at those 3 numbers is worth less than nothing, because the first person to check it will know it is wrong.
One construction helps rather than just measuring. Rather than one width everywhere, the range can be scaled by how difficult the model finds each case, tightening on a fresh tool by 21% and roughly doubling on a worn one. Coverage on worn tools rises from 68.0% to 78.6%. Whether that reaches the promised 90% this data cannot say, since the interval on it runs from 57.4 to 94.4. What can be said is narrower and still useful: it is the best of the constructions tested there, and the only one that widens where the model is weak instead of pretending otherwise.
- Pair each tool change with what was actually measured on the tool, not with why it was changed.
- Decide the repeated unit before splitting anything: insert, batch, machine, programme.
- Score whatever range you quote on units the model has not seen.
- Report the result by wear regime rather than in aggregate, because the aggregate is dominated by the easy first half of the tool’s life.
- Convert to cuts thrown away and cuts run long before converting to money.
Assumption. That a policy acts at the first cut meeting its condition. Real shops act at shift ends and batch boundaries, which moves every number in the table and does not change their order.
What This Can and Cannot Show
Sixteen inserts, cast iron and stainless, on a machine older than most of the people running one. The mechanism transfers. The numbers do not.
8 limits bound what can be carried out of this article.
- The materials are cast iron and stainless J45, on a machining centre of 1990s vintage. Nothing here transfers numerically to titanium, to nickel alloys, or to a modern 5-axis cell. The mechanism does.
- 16 inserts is a small number of units. Every interval in Fig. 02 comes from resampling whole inserts rather than rows, because 146 cuts from 16 inserts are not 146 independent trials. Each resample draws 16 inserts with replacement and takes every cut belonging to them, so an unusual tool is either in or out rather than diluted. Those intervals are wide, and they are the honest ones.
- The worn-tool results rest on fewer inserts than they appear to. 20 observations, from 10 distinct inserts, and 6 of the 20 come from a single insert, the one that ran to 1.53 mm. One tool carries about 30% of the evidence for the most quotable claim in the article.
- The 0.60 mm change point is conventional, not a property of the data. Every result was recomputed from 0.30 to 1.00 mm, and the finding strengthens as the limit rises. The width claim in § 04 is bounded at 0.65 mm for that reason and is stated bounded.
- The model cannot reach the top of the wear range, as § 05 describes. Thresholds above 0.73 mm never fire.
- Wear was measured off the machine, by microscope, at irregular intervals. How much the measurement itself varies is undocumented and appears in none of these numbers.
- The policy simulation is not a deployment. Wear is known only at the cuts where it was measured, and policies act immediately rather than at a shift boundary.
- Nothing here says a monitoring system will not help. It says this dataset is too small to show that it does.
3 progressively richer sets of sensor features were built from the 6 channels, and every one made prediction worse: 0.165 mm from cut count alone, against 0.176, 0.193 and 0.194 mm as channels were added.
The honest reading is about sample size and not about sensors. With 16 inserts and roughly 110 training rows, a model with 58 inputs spends more in noise than it recovers in signal. This is evidence that this dataset cannot justify a monitoring investment. It is not evidence that monitoring does not work, and anyone quoting it that way is quoting it wrongly.
The useful inversion for a shop weighing one: if your entire history is 16 tool lives, you do not yet have the data to justify the sensor package, and an audit will tell you that before you buy it rather than after.
What none of this can show is whether your process behaves like this one. Your alloys are not these, and neither is your change rule.
If you buy a tool monitoring system that quotes confidence, everything above applies to it exactly as it applies to the model here. Coverage is checkable by the customer, from records the customer already holds, without the vendor’s cooperation and without their model. Recipe 5.1 doubles as the questions to put in the next review.
Tool Wear and Process Calibration Audit: the analysis above, run on your records rather than on a public benchmark. Input: wear or dimensional records, cutting parameters, intervention history and scrap logs; 6 months or more, exported however they already exist, including a spreadsheet. You get: the coverage picture for your own process, by regime rather than in aggregate; the change-threshold curve at your own limit with your own rates; a data structure check delivered on day 2, not day 14, because if the records cannot support the analysis both of us should know inside 48 hours. Terms: fixed price, two weeks from data receipt, standard mutual NDA before anything moves, your data deleted on delivery and never used elsewhere. If your current change interval is already well set, the audit says so, which is a result worth having too.
Matthew Tanti · Kwantil, Malta · contact via kwantil.com
@article{tanti2026,
author = {Matthew Tanti},
title = {Tool Change Intervals, Audited Against Measured Wear},
journal = {Kwantil},
year = {2026},
number = {KW-2026-03},
version = {1.0},
url = {https://kwantil.com/papers/tool-wear-calibration-audit/}
}Matthew Tanti (2026). Tool Change Intervals, Audited Against Measured Wear. Kwantil, KW-2026-03, v1.0.