Premium · 3 plants, 36 CNC machines, six months

SPC for Precision Machining

Six months of control charts with the ground truth underneath

504,786 rows in 15 linked tables: 3 plants, 36 CNC machines, 158 controlled characteristics, 82,950 subgroups and 414,750 individual gauge readings from 2025-01-06 to 2025-06-29, with every cause of every shift recorded next to the data.

Why it exists

Real SPC data never tells you what actually happened. Here, every subgroup carries the true size of the process shift that produced it and where it came from: an assignable cause on the machine, tool wear since the last change, the material lot, the operator's setup, or the process's own drift. The control charts are computed exactly the way a quality engineer would compute them, so you can measure what the charts caught, what they missed, and how fast.

Uses

What it is built for

  • Control charts and capability

    X-bar/R and X-bar/s, Phase I limits from the first 25 subgroups, Cp, Cpk, Pp, Ppk and observed ppm per characteristic, two-sided and upper-limit-only (flatness, true position, surface roughness).

  • Run-rule and anomaly detection with labels

    the six Western Electric/Nelson flags are in the file; the true shift is too, so detection rate, false-alarm rate and average run length can be measured rather than assumed. A benchmark for ML detectors (CUSUM, EWMA, isolation forests, sequence models) against a known truth.

  • Root cause analysis

    join a signal to the cause active on its machine, the material lot, the operator, the hours since the last tool change and the days since maintenance.

  • Tool wear and predictive maintenance

    shaft diameters and surface roughness grow and bores shrink with the hours on a tool, and reset at each change; within-subgroup spread widens as a machine goes longer without maintenance.

  • Measurement system analysis

    each gauge has a crossed GR&R study (ANOVA or range method) and a calibration history; readings sit on the gauge's own resolution, so a coarse gauge shows up as it does on a real chart.

  • Teaching and dashboards

    realistic enough for an SPC course, a Six Sigma project, a Power BI/Tableau quality dashboard, or a MES/QMS demo database.

In the data

What the full files show

Computed from every row of the full dataset when it was packaged, not drawn as a target. The free preview is a slice of the same data.

An X-bar chart with the truth underneath: overall length, characteristic 139

109.6615109.6808109.7000UCLCLLCLJan 2025Feb 2025Mar 2025Apr 2025May 2025Jun 2025
  • Subgroup mean
  • True assignable cause in force
  • Rule tripped
Every subgroup mean for one characteristic over the six months (mm), with its Phase I centre line and control limits. Shaded spans are when an assignable cause was really in force (assignable_causes); dots are subgroups that trip a Western Electric rule. Some causes trip nothing, some alarms have no cause: both are labelled in the data.

How often the rules fire, by the size of the true shift

  • no cause20.1%
  • under 0.5 sigma38.1%
  • 0.5 to 1 sigma48.9%
  • 1 to 1.5 sigma72.6%
  • 1.5 to 2 sigma89.5%
  • 2 or more sigma97.8%
Share of Phase II subgroups that trip at least one rule, grouped by the shift of the assignable cause in force (true_cause_shift_sigma, in within-subgroup sigma). Small shifts mostly slip through; tool wear, lots and drift fire some rules even with no cause.

Assignable causes, by category

  • tool wear beyond limit79
  • thermal growth58
  • fixture loosened43
  • material hardness out of range37
  • coolant concentration32
  • wrong offset entered30
  • spindle bearing wear21
Every special cause in assignable_causes, with its machine, start, end, size and the corrective action taken.

Answer key

The truth, in its own columns

What a model, a control chart or an analyst is trying to find ships next to the data, so you can score an answer instead of guessing at it. Each line is measured from the full files.

  • subgroups.active_cause_idThe assignable cause in force when the subgroup was sampled, if any.285 distinct, 95% empty
  • subgroups.true_cause_shift_sigmaThat cause's shift of the mean, in within-subgroup sigma.-4.36 to 4.62, mean 0.021
  • subgroups.true_tool_wear_sigmaTool wear since the last change, in sigma.-0.6 to 0.6, mean 0.03
  • subgroups.true_lot_shift_sigmaThe material lot's hardness effect, in sigma.-0.17 to 0.29, mean 0.0099
  • subgroups.true_operator_shift_sigmaThe operator's setup bias, in sigma.-0.32 to 0.21, mean 0.00061
  • subgroups.true_driftThe process mean's own slow drift (AR(1)), in sigma.-0.3 to 0.32, mean 0.00016
  • subgroups.true_process_meanThe true process mean at that moment, in the characteristic's unit.-0.038 to 156.5, mean 32.57
  • subgroups.noise_factorHow much wider the within-subgroup spread is as maintenance ages.1 to 1.2, mean 1.06
  • assignable_causes.magnitude_sigmaEach cause's size, with its category, start, end and direction.0.45 to 4.62, mean 1.55

How it behaves

Measured on the files you download

Each of these is computed from the delivered rows when the dataset is packaged, not written as a target.

  • 5.0% of subgroups were sampled while an assignable cause was active; 94% of the Phase II subgroups under a cause of 1.5 sigma or more trip at least one rule, against 20% of those under none (tool wear, lots and drift still move the mean there, and are labelled).
  • Rule 1 false-alarm rate, stable Phase II: 0.78% (theory 0.27%); estimated limits from 25 subgroups and an autocorrelated process put it above the textbook 0.27%, as they do in practice.
  • 0.28% of readings are out of specification. 9% of characteristics have Cpk below 1.0 and 65% are at 1.33 or better; 25% are marked critical to quality.

Audit

78 of 78 checks pass

Re-run on the delivered files by an independent script with plain pandas. The results ship in INTEGRITY.json.

  • every key resolves and every primary key is unique; each characteristic's machine and gauge are on its part's line, each operator on the subgroup's line, each lot belongs to the part and was received before it was used;
  • every subgroup has exactly 5 readings; X-bar, max, min, range and s re-add exactly from the readings;
  • control limits, Cpk and Ppk recompute from the subgroups and readings; all six rule flags recompute with 0 differences;
  • each feature is measured on a gauge that suits it, and every reading sits on its gauge's resolution;
  • a cause is only recorded on a subgroup sampled inside that cause's window on that machine;
  • the gauge R&R trials recover each gauge's stated %GR&R.
All 78 checks
  • ✓ plants.plant_id unique
  • ✓ production_lines.line_id unique
  • ✓ machines.machine_id unique
  • ✓ gauges.gauge_id unique
  • ✓ operators.operator_id unique
  • ✓ parts.part_id unique
  • ✓ characteristics.characteristic_id unique
  • ✓ material_lots.lot_id unique
  • ✓ tool_changes.tool_change_id unique
  • ✓ maintenance_events.maintenance_id unique
  • ✓ assignable_causes.cause_id unique
  • ✓ subgroups.subgroup_id unique
  • ✓ readings.reading_id unique
  • ✓ gauge_calibrations.calibration_id unique
  • ✓ gauge_rr_trials.trial_id unique
  • ✓ production_lines.plant_id -> plants: 0 orphans of 9
  • ✓ machines.line_id -> production_lines: 0 orphans of 36
  • ✓ gauges.line_id -> production_lines: 0 orphans of 36
  • ✓ operators.line_id -> production_lines: 0 orphans of 72
  • ✓ parts.line_id -> production_lines: 0 orphans of 40
  • ✓ characteristics.part_id -> parts: 0 orphans of 158
  • ✓ characteristics.machine_id -> machines: 0 orphans of 158
  • ✓ characteristics.gauge_id -> gauges: 0 orphans of 158
  • ✓ material_lots.part_id -> parts: 0 orphans of 536
  • ✓ tool_changes.machine_id -> machines: 0 orphans of 2200
  • ✓ maintenance_events.machine_id -> machines: 0 orphans of 240
  • ✓ assignable_causes.machine_id -> machines: 0 orphans of 300
  • ✓ subgroups.characteristic_id -> characteristics: 0 orphans of 82950
  • ✓ subgroups.operator_id -> operators: 0 orphans of 82950
  • ✓ subgroups.material_lot_id -> material_lots: 0 orphans of 78215
  • ✓ subgroups.active_cause_id -> assignable_causes: 0 orphans of 4140
  • ✓ readings.subgroup_id -> subgroups: 0 orphans of 414750
  • ✓ gauge_calibrations.gauge_id -> gauges: 0 orphans of 216
  • ✓ gauge_rr_trials.gauge_id -> gauges: 0 orphans of 3240
  • ✓ characteristic line = its part's line
  • ✓ characteristic machine on its line
  • ✓ characteristic gauge on its line
  • ✓ subgroup operator on its line
  • ✓ every line has one gauge of each kind
  • ✓ each feature measured on a gauge that suits it
  • ✓ readings sit on their gauge's resolution: max off-grid 0.0000 steps
  • ✓ subgroup lot belongs to its part
  • ✓ subgroup lot received before sampling
  • ✓ subgroups with a lot: 94.3%
  • ✓ 5 readings per subgroup
  • ✓ sample_no 1..5
  • ✓ subgroups.x_bar re-adds from readings: 0 differ
  • ✓ subgroups.x_max re-adds from readings: 0 differ
  • ✓ subgroups.x_min re-adds from readings: 0 differ
  • ✓ subgroups.s re-adds from readings: 0 differ
  • ✓ x_range = x_max - x_min
  • ✓ measured_at = sampled_at
  • ✓ sample_seq is time order
  • ✓ control limits = Phase I X-bar/R (A2 .577, D4 2.114)
  • ✓ Cpk recomputes (R-bar/d2): max diff 0.0024
  • ✓ Ppk recomputes (overall sd): max diff 0.0013
  • ✓ Ppk <= Cpk mostly (shifts and drift add variation): 95%
  • ✓ rule1_beyond_3sigma recomputes: 0 differ; fires on 4.4%
  • ✓ rule2_2of3_beyond_2sigma recomputes: 0 differ; fires on 7.8%
  • ✓ rule3_4of5_beyond_1sigma recomputes: 0 differ; fires on 11.2%
  • ✓ rule4_8_same_side recomputes: 0 differ; fires on 9.5%
  • ✓ trend_6_points recomputes: 0 differ; fires on 0.4%
  • ✓ range_beyond_ucl recomputes: 0 differ; fires on 1.1%
  • ✓ out_of_control = any rule
  • ✓ stable process rarely signals (rule 1): 0.78%
  • ✓ rule 1 false-alarm rate, stable Phase II: 0.78% (theory 0.27%); OOC while a >=1.5 sigma cause is active: 94%
  • ✓ causes are what the charts catch: 94% vs 20% without
  • ✓ cause active only inside its window
  • ✓ share of subgroups under a cause: 5.0%
  • ✓ x-bar is autocorrelated (drift, wear, lots), not white noise: mean lag-1 autocorrelation 0.33
  • ✓ true shifts of 3 sigma-xbar or more are signalled: 94% of 1912
  • ✓ OD grows / bore shrinks with hours since tool change: corr 0.24
  • ✓ out-of-spec readings rare but present: 0.28% (1147 readings)
  • ✓ Cpk spread: some incapable (<1), most capable (>1.33): <1: 9%, >1.33: 65%
  • ✓ upper-only features have no LSL
  • ✓ no negative flatness/position/roughness
  • ✓ GR&R: 10 parts x 3 appraisers x 3 trials
  • ✓ GR&R repeatability tracks gauge %GRR: corr 0.96
  • ✓ GR&R part effect dominates (parts differ)

Tables

15 tables, 504,786 rows

Every table in the zip with what it holds, its rows and its columns. The bars are on a log scale, so the small reference tables still show.

  • readingsEvery individual measurement, on its gauge's resolution, with out-of-spec flag414,750 rows · 8 cols
  • subgroupsOne row per n=5 sample: X-bar, R, s, Western Electric flags, the lot, operator, hours since tool change, and the TRUE cause of any shift82,950 rows · 36 cols
  • gauge_rr_trialsA crossed gauge R&R study per gauge: 10 parts x 3 appraisers x 3 trials3,240 rows · 7 cols
  • tool_changesEvery insert, drill or wheel change and why2,200 rows · 5 cols
  • material_lotsBar-stock heats per part: supplier, received date, hardness536 rows · 6 cols
  • assignable_causesEvery special cause: machine, category, start, end, size in sigma, direction, corrective action, CAPA300 rows · 10 cols
  • maintenance_eventsPreventive and corrective maintenance with downtime240 rows · 5 cols
  • gauge_calibrationsCalibration history with as-found error and result216 rows · 7 cols
  • characteristicsThe drawing's controlled features: spec limits, gauge, Phase I control limits, Cp/Cpk, Pp/Ppk, observed ppm158 rows · 40 cols
  • operatorsOperators by line and home shift, experience, certification, setup bias72 rows · 7 cols
  • partsMachined parts: family, material, sector, drawing revision, annual volume40 rows · 8 cols
  • gaugesOne CMM, air bore gauge, micrometer and profilometer per line: resolution, %GR&R, repeatability and reproducibility36 rows · 10 cols
  • machinesCNC machines, each with its own capability factor and PM interval36 rows · 7 cols
  • production_linesMachining lines and their process family9 rows · 5 cols
  • plantsThree plants and their region3 rows · 3 cols

Explore

Every table, profiled

Each column's type, spread, empties and most common values, measured from the full CSVs. Switch to the first rows to see the data as it sits in the file.

readings.csv

414,750 rows · 8 columns · 2 foreign keys

reading_idprimary key
unique on every row
414,750 distinctno empties
subgroup_idforeign key
points to subgroups.subgroup_id
82,950 distinctno empties
characteristic_idforeign key
points to characteristics.characteristic_id
158 distinctno empties
sample_nonumber
15
mean 3median 31 to 5no empties
measured_atdate
Jan 2025Jun 2025
2025-01-06 to 2025-06-29no empties
valuenumber
0.004156.5
mean 32.57median 21.620 to 156.5no empties
deviation_from_targetnumber
-0.560.78
mean 0.012median 0.0002-1.38 to 2.54no empties
is_out_of_specyes / no
true · 0.3%false · 100%
no empties

Keys

Every join resolves

24 foreign keys, 1,334,162 references checked against the table each one points at. None points at a row that does not exist.

assignable_causes

  • machine_idmachines.machine_id0 orphans

characteristics

  • part_idparts.part_id0 orphans
  • line_idproduction_lines.line_id0 orphans
  • machine_idmachines.machine_id0 orphans
  • gauge_idgauges.gauge_id0 orphans

gauge_calibrations

  • gauge_idgauges.gauge_id0 orphans

gauge_rr_trials

  • gauge_idgauges.gauge_id0 orphans

gauges

  • line_idproduction_lines.line_id0 orphans

machines

  • line_idproduction_lines.line_id0 orphans

maintenance_events

  • machine_idmachines.machine_id0 orphans

material_lots

  • part_idparts.part_id0 orphans

operators

  • line_idproduction_lines.line_id0 orphans

parts

  • line_idproduction_lines.line_id0 orphans

production_lines

  • plant_idplants.plant_id0 orphans

readings

  • subgroup_idsubgroups.subgroup_id0 orphans
  • characteristic_idcharacteristics.characteristic_id0 orphans

subgroups

  • characteristic_idcharacteristics.characteristic_id0 orphans
  • part_idparts.part_id0 orphans
  • machine_idmachines.machine_id0 orphans
  • line_idproduction_lines.line_id0 orphans
  • operator_idoperators.operator_id0 orphans
  • material_lot_idmaterial_lots.lot_id0 orphans
  • active_cause_idassignable_causes.cause_id0 orphans

tool_changes

  • machine_idmachines.machine_id0 orphans

In the zip

What you get

  • CSVs: one file per table, with a header row and ISO dates.
  • README.md: what each table holds, how the data behaves (measured), what was checked and what to know.
  • INTEGRITY.json: the 78 audit checks and their results.
  • RECIPE.json: the Misata blueprint that made these exact rows, seed 20250106.
All 18 files
  • plants.csv93 B
  • production_lines.csv278 B
  • machines.csv2 KB
  • gauges.csv3 KB
  • operators.csv3 KB
  • parts.csv3 KB
  • characteristics.csv42 KB
  • material_lots.csv26 KB
  • tool_changes.csv124 KB
  • maintenance_events.csv10 KB
  • assignable_causes.csv32 KB
  • subgroups.csv17.7 MB
  • readings.csv24.2 MB
  • gauge_calibrations.csv15 KB
  • gauge_rr_trials.csv86 KB
  • README.md7 KB
  • INTEGRITY.json11 KB
  • RECIPE.json37 KB

Before you use it

Things to know

  • Each characteristic is sampled about three times a day (525 subgroups over the six months) at the machine's own times, not on an exact clock.
  • A rule flag is true on every subgroup where its condition holds, not only on the first: eight points on one side flag the eighth, ninth and so on.
  • design_cpk is what the process was designed for; cpk is what the Phase I data shows.
  • A subgroup sampled before its part's first lot arrived in the window has no lot id (the lot in use came before the window); days_since_maintenance and hours_since_tool_change are empty before a machine's first event in the window.
  • Plants, parts, suppliers and people are invented; no real company, product or person is represented.

Questions

Before you buy

Is this real data?
No. This is synthetic data generated by software. No row describes a real person, company, patient, store, machine or transaction. The patterns are modelled to be realistic and the statistics quoted are measured on these files, but they do not describe any real population or market. Use it for learning, testing, demos, benchmarks and prototyping, not as evidence about the real world. Provided as is, without warranty.
What do I get?
One zip of 8.9 MB: 15 linked tables and 504,786 rows as CSV, a README of what each table holds and how the data behaves, INTEGRITY.json with the 78 audit checks, and RECIPE.json, the Misata blueprint that made these exact rows (seed 20250106).
Can I try it before buying?
Yes. The free preview (110 KB) is a slice of the same data with the keys intact, plus the README, so you can load it and check it fits before you pay.
How was it checked?
78 of 78 checks pass, re-run on the delivered files by an independent script with plain pandas: every key resolves and every primary key is unique; each characteristic's machine and gauge are on its part's line, each operator on the subgroup's line, each lot belongs to the part and was received before it was used; every subgroup has exactly 5 readings; X-bar, max, min, range and s re-add exactly from the readings; control limits, Cpk and Ppk recompute from the subgroups and readings; all six rule flags recompute with 0 differences; each feature is measured on a gauge that suits it, and every reading sits on its gauge's resolution; a cause is only recorded on a subgroup sampled inside that cause's window on that machine; the gauge R&R trials recover each gauge's stated %GR&R.
Can I use it commercially?
Yes: in any project, course, benchmark, demo or product, commercial or not. You may not resell or redistribute the dataset itself as a dataset.
What should I know before using it?
Each characteristic is sampled about three times a day (525 subgroups over the six months) at the machine's own times, not on an exact clock. A rule flag is true on every subgroup where its condition holds, not only on the first: eight points on one side flag the eighth, ninth and so on. design_cpk is what the process was designed for; cpk is what the Phase I data shows. A subgroup sampled before its part's first lot arrived in the window has no lot id (the lot in use came before the window); days_since_maintenance and hours_since_tool_change are empty before a machine's first event in the window. Plants, parts, suppliers and people are invented; no real company, product or person is represented.
Can I get a bigger or different version?
Yes. RECIPE.json runs in Misata Studio or through Misata's MCP server to make a variant, or ask us to build one to your spec.

Want it bigger, in another setting, or with your own columns? Have us build it or make it in Studio.

More premium datasets

All 4
  • readings, rows per monthJan 2024 – Dec 2025

    240 machines at 4 plants, two years

    Two years of sensor readings, failures, repairs and costs for 240 machines

    Rows
    179,029
    Tables
    6
    Checks
    62 of 62
    Profile and preview
  • fact_store_item_day, rows per monthFeb 2024 – Jan 2026

    Grocery chain, 10 stores, fiscal 2024 and 2025

    Two fiscal years of a grocery chain, from the shelf to the receipt

    Rows
    2,780,138
    Tables
    14
    Checks
    110 of 110
    Profile and preview
  • claims, rows per monthJan 2025 – Dec 2025

    10 payers, 8 facilities, calendar 2025

    A year of revenue cycle with the reason behind every denial

    Rows
    356,609
    Tables
    8
    Checks
    87 of 87
    Profile and preview