Premium · 240 machines at 4 plants, two years
Predictive Maintenance Fleet
Two years of sensor readings, failures, repairs and costs for 240 machines
179,029 rows in 6 linked tables: 240 rotating assets at 4 plants, 175,440 daily sensor readings, 771 asset lives (runs), 532 work orders and 2,042 route inspections, 1 January 2024 to 31 December 2025.
Real run-to-failure data is scarce: plants replace parts before they fail, failures are rare, and the few public sets are one machine on a test rig. This fleet has the whole picture a reliability team works with: many machine types and failure modes, sensors that respond to the mode that is developing, repairs that cut lives short (censoring), and the money. Every reading carries the true remaining life, even where the part was replaced before it could fail, so a remaining-life model can be scored against what would really have happened.
Uses
What it is built for
Remaining useful life (RUL) regression
with
rul_daysortrue_rul_days, and failure within 30 days classification withfails_within_30d(5.6% positive, the realistic imbalance).Survival analysis
239 runs are right-censored at the cut-off and 204 were ended early by a repair or replacement; Weibull, Cox and Kaplan-Meier on
runs.csvwith load, climate and failure mode as covariates.Anomaly detection and alarm tuning
the alarms are rule-based (ISO zone, temperature, pressure, current, debris); beat them with a model, and measure false alarms against lead time.
Failure mode diagnosis
from the sensor signature: vibration for bearings, imbalance and misalignment, temperature for lubrication, pressure for impellers and seals, current for windings, oil debris for gears.
Maintenance strategy and cost
run-to-failure against condition-based against age-based replacement; breakdowns cost a median $30,991 against $9,796 for a condition-based repair.
Dashboards and teaching
MTBF, MTTR, availability, backlog, PM compliance, cost by site and criticality in Power BI, Tableau or SQL; a reliability engineering course or a CMMS demo.
In the data
What the full files show
Computed from every row of the full dataset when it was packaged, not drawn as a target. The free preview is a slice of the same data.
Vibration through a life that ends in failure
- Gave warning
- Sudden
Chance of failing within 30 days, by ISO 10816-3 zone
- Zone A2.3%
- Zone B8.3%
- Zone C26%
- Zone D85.3%
Failures by mode
- bearing wear144
- shaft misalignment81
- seal leak72
- rotor imbalance57
- lubrication breakdown47
- impeller erosion36
- gear tooth wear34
- winding insulation28
What a work order costs, by type
- corrective (breakdown)$30,991
- corrective (planned, condition-based)$9,796
- preventive replacement$8,348
Answer key
The truth, in its own columns
What a model, a control chart or an analyst is trying to find ships next to the data, so you can score an answer instead of guessing at it. Each line is measured from the full files.
- runs.failure_modeThe true failure mode of the life, which sets the signals that move.8 values: bearing wear, shaft misalignment, seal leak…
- runs.true_life_daysHow long the life lasts to failure if nothing intervenes; a run caught or replaced early ends before it.52.8 to 3,979.8, mean 428.9
- runs.true_failure_atWhen the machine would fail, including for runs that were caught or renewed first.2024-01-11 19:12:00 to 2035-08-01 22:00:00
- runs.warning_signsWhether the failure gives signs a sensor can see, or comes suddenly.true on 82.2% of rows
- runs.pf_onsetThe share of the life at which the signals start to rise (the P on the P-F curve); near 1 for sudden failures.0.55 to 0.98, mean 0.69
- readings.life_usedThe share of the run's true life used at the reading.0 to 1, mean 0.46, 0.78% empty
- readings.rul_daysDays left until the run ends, however it ends.0 to 3,979.4, mean 289.6, 0.78% empty
- readings.true_rul_daysDays left until the true failure: the remaining-useful-life target.0 to 3,979.4, mean 325.6, 0.78% empty
- readings.fails_within_30dWhether the machine truly fails within 30 days: the classification target.true on 5.6% of rows
How it behaves
Measured on the files you download
Each of these is computed from the delivered rows when the dataset is packaged, not written as a target.
- Signals stay flat for most of a life, then rise along a P-F curve: median vibration on vibration-type failures goes 1.62 -> 1.63 -> 2.50 -> 8.62 mm/s from new to the last 3% of life.
- 91% of warned and 7% of sudden failures gave a sustained alarm first, median 46 days ahead. 25% of failures were sudden: some faults give no warning a sensor can see.
- Harder-worked machines live shorter (bearings with the cube of load): Spearman -0.57 between load factor and true life; hot sites wear machines out sooner (median true life 275 against 333 days).
- Near the end of a life, bearing temperature is 62 C against 42 C for bearing and lubrication failures against other modes; discharge pressure falls to 78% of nominal against 96% for impeller and seal failures.
- Breakdowns keep a machine down a median 51 h against 6 h for planned work; lost production is $20,565,604 of the $22,049,444 maintenance cost.
- Runs: 328 failures, 123 caught by monitoring, 81 age-based replacements, 239 still running at the cut-off. Sensor dropouts: about 0.8% of readings.
Audit
62 of 62 checks pass
Re-run on the delivered files by an independent script with plain pandas. The results ship in INTEGRITY.json.
- every key resolves; one reading per asset per day for all 731 days; runs of an asset never overlap and each starts after the last repair ended;
- the run in force is rebuilt for every reading from the run table, and every label (run age, life used, remaining life, true remaining life, fails within 30 days) recomputes from it;
- no sensor value while the asset is down; pressure only on pumps and compressors, oil debris only on gearboxes and compressors; ISO 10816-3 zone, alarm and alarm reason recompute from the readings;
- one work order per run ended in the window, its type from how the run ended, its times from the run, its costs re-adding; inspection findings recompute from the measured vibration and follow each asset's PM interval;
- the physics above: P-F curve, mode-specific signals, load and climate against life, criticality against catches, sudden failures, false alarm rate.
All 62 checks
- ✓ sites.site_id unique
- ✓ assets.asset_id unique
- ✓ runs.run_id unique
- ✓ readings.reading_id unique
- ✓ work_orders.work_order_id unique
- ✓ inspections.inspection_id unique
- ✓ assets.site_id -> sites: 0 orphans of 240
- ✓ runs.asset_id -> assets: 0 orphans of 771
- ✓ runs.site_id -> sites: 0 orphans of 771
- ✓ readings.asset_id -> assets: 0 orphans of 175,440
- ✓ readings.site_id -> sites: 0 orphans of 175,440
- ✓ readings.run_id -> runs: 0 orphans of 174,063
- ✓ work_orders.run_id -> runs: 0 orphans of 532
- ✓ work_orders.asset_id -> assets: 0 orphans of 532
- ✓ inspections.asset_id -> assets: 0 orphans of 2,042
- ✓ a run's site and a reading's site are the asset's
- ✓ one reading per asset per day, every day of both years: 731 days x 240 assets
- ✓ a run's end = start + life; true failure = start + true life; back in service = end + downtime
- ✓ a failure ends at the true life; a renewal before it
- ✓ each run starts the moment the last repair ends
- ✓ run numbers count up from 1 per asset
- ✓ each asset is part-way through a run on the first day
- ✓ no run starts after the cut-off
- ✓ outcome at cut-off: running when the run outlives the data
- ✓ every failure mode can happen to its kind of asset
- ✓ run_id is the run in force at the reading, empty while the asset is down
- ✓ running exactly when a run is in force; no run hours while down
- ✓ rul_days = run end - reading
- ✓ true_rul_days = true failure - reading (past a renewal too)
- ✓ life_used = age / true life
- ✓ fails_within_30d only in a run's last 30 days before a real failure
- ✓ no sensor value while the asset is down
- ✓ sensor dropouts are rare (about 1%): vibration_rms_mm_s 0.84%, bearing_temp_c 0.78%, motor_current_a 0.80%
- ✓ pressure only on pumps and compressors, oil debris only on gearboxes and compressors
- ✓ run hours are the asset's duty hours while running
- ✓ ISO 10816-3 zone recomputes from vibration
- ✓ alarm reason recomputes (vibration zone C/D, bearing over 85 C, pressure under 88% or current over 108% of expected at that load, oil debris over 1500/ml)
- ✓ daily load wanders around each asset's own load factor: median gap 0.5 points
- ✓ readings stay physical
- ✓ one work order for every run that ended inside the window, none for others
- ✓ work order type follows how the run ended
- ✓ work starts when the run ends and closes when the asset is back
- ✓ a breakdown is raised when it happens, planned work before
- ✓ failure mode and component replaced are the run's
- ✓ total cost = parts + labour + lost production
- ✓ labour cost = hours x rate (breakdown overtime 135, planned 95)
- ✓ found condition: failed on breakdowns, a confirmed defect on monitored catches
- ✓ an inspection finds the asset down exactly when no run is in force
- ✓ inspection finding recomputes from the measured vibration
- ✓ route frequency follows each asset's PM interval: 6-12 per asset
- ✓ vibration flat for most of the life, then rising toward a vibration failure: 1.62 -> 1.63 -> 2.50 -> 8.62
- ✓ bearings and lubrication run hot near the end; other modes do not: 62 C against 42 C
- ✓ impeller erosion and seal leaks lose pressure: 78% of nominal against 96%
- ✓ hot sites wear equipment out sooner: median true life 275 against 333 days
- ✓ harder-worked assets live shorter: Spearman -0.57
- ✓ critical assets have more failures caught in time: A 49%, B 19%, C 8%
- ✓ breakdowns keep an asset down far longer than planned work: 51 h against 6 h
- ✓ failures with warning signs alarm first, weeks ahead (the P-F interval); sudden ones mostly do not: 91% of warned and 7% of sudden failures gave a sustained alarm first, median 46 days ahead; 25% of failures sudden
- ✓ only a run with warning signs is caught by monitoring
- ✓ false alarms happen on healthy machines, rarely: 0.17% of readings in the first half of a life
- ✓ zone D is rare: 1.9% of readings
- ✓ every asset ends the data running or under repair; runs are right-censored at the cut-off: 239 censored runs
- ✓ runs: 771; failure 328, running at cut-off 239, caught by monitoring 123, age-based replacement 81
- ✓ failure modes: bearing wear 32%, shaft misalignment 15%, seal leak 12%, lubrication breakdown 11%, rotor imbalance 11%, impeller erosion 7%, gear tooth wear 6%, winding insulation 6%
- ✓ maintenance spend: $22,049,444, of it lost production $20,565,604
Tables
6 tables, 179,029 rows
Every table in the zip with what it holds, its rows and its columns. The bars are on a log scale, so the small reference tables still show.
- readingsOne reading per asset per day for two years: vibration RMS and crest factor, bearing temperature, motor current, discharge pressure, oil debris, load, ambient, ISO 10816-3 zone, alarm and its reason; plus labels (run age, life used, remaining life to the run's end and to the true failure, fails within 30 days)175,440 rows · 24 cols
- inspectionsRoute-based vibration checks at each asset's PM interval, with the measured value and the finding2,042 rows · 7 cols
- runsEvery life of every asset, from install or repair to the next failure or renewal: failure mode, true life, whether it gave warning signs, how it ended, downtime, and whether it was still running at the cut-off771 rows · 19 cols
- work_ordersOne per run that ended in the window: breakdown, condition-based or preventive, priority, raised/started/completed, component replaced, found condition, labour hours, parts, labour and lost-production cost532 rows · 19 cols
- assetsPumps, motors, fans, gearboxes and compressors: type, criticality A/B/C, rated power and current, nominal pressure, install year, load factor, duty hours, online monitoring or route-based, PM interval240 rows · 15 cols
- sitesFour plants in different climates: ambient temperature and its seasonal swing, spares lead time, cost of an hour of downtime4 rows · 7 cols
Explore
Every table, profiled
Each column's type, spread, empties and most common values, measured from the full CSVs. Switch to the first rows to see the data as it sits in the file.
6 tables
readings.csv
175,440 rows · 24 columns · 3 foreign keys
- A79%
- B15%
- C3.8%
- D1.9%
- vibration82%
- pressure loss14%
- oil debris2.2%
- motor current2.1%
- bearing temperature0.1%
Keys
Every join resolves
10 foreign keys, 530,363 references checked against the table each one points at. None points at a row that does not exist.
assets
- site_idsites.site_id240 rows checked0 orphans
inspections
- asset_idassets.asset_id2,042 rows checked0 orphans
readings
- asset_idassets.asset_id175,440 rows checked0 orphans
- site_idsites.site_id175,440 rows checked0 orphans
- run_idruns.run_id174,063 rows checked0 orphans
runs
- asset_idassets.asset_id771 rows checked0 orphans
- site_idsites.site_id771 rows checked0 orphans
work_orders
- run_idruns.run_id532 rows checked0 orphans
- asset_idassets.asset_id532 rows checked0 orphans
- site_idsites.site_id532 rows checked0 orphans
In the zip
What you get
- CSVs: one file per table, with a header row and ISO dates.
- README.md: what each table holds, how the data behaves (measured), what was checked and what to know.
- INTEGRITY.json: the 62 audit checks and their results.
- RECIPE.json: the Misata blueprint that made these exact rows, seed 20240101.
All 9 files
- sites.csv222 B
- assets.csv20 KB
- runs.csv145 KB
- readings.csv23.2 MB
- work_orders.csv109 KB
- inspections.csv184 KB
- README.md8 KB
- INTEGRITY.json9 KB
- RECIPE.json27 KB
Before you use it
Things to know
- ISO 10816-3 zone limits (2.3 / 4.5 / 7.1 mm/s) are applied to every asset for simplicity; real limits depend on machine class and foundation.
true_rul_daysandtrue_failure_atare ground truth no plant could know; userul_daysfor a realistic label and the true values to study how censoring biases a model.pf_onsetis the fraction of the true life at which degradation became visible (about 0.985 for a sudden failure).- Lives are compressed so two years hold enough failures to learn from: the median life of a failure mode is 260 to 760 days, several times shorter than most plants see (a well-kept bearing or gear set often runs for years). The shape of each life (the P-F curve, the load and climate effects, the censoring) is what the data is built to teach, not its length; do not quote the MTBF as a benchmark.
lost_production_costis downtime hours times the site's hourly rate, scaled by criticality (A 1, B 0.45, C 0.12).- Sites, assets and models are invented; no real plant, manufacturer or equipment record is represented.
Questions
Before you buy
- Is this real data?
- No. This is synthetic data generated by software. No row describes a real person, company, patient, store, machine or transaction. The patterns are modelled to be realistic and the statistics quoted are measured on these files, but they do not describe any real population or market. Use it for learning, testing, demos, benchmarks and prototyping, not as evidence about the real world. Provided as is, without warranty.
- What do I get?
- One zip of 6.1 MB: 6 linked tables and 179,029 rows as CSV, a README of what each table holds and how the data behaves, INTEGRITY.json with the 62 audit checks, and RECIPE.json, the Misata blueprint that made these exact rows (seed 20240101).
- Can I try it before buying?
- Yes. The free preview (303 KB) is a slice of the same data with the keys intact, plus the README, so you can load it and check it fits before you pay.
- How was it checked?
- 62 of 62 checks pass, re-run on the delivered files by an independent script with plain pandas: every key resolves; one reading per asset per day for all 731 days; runs of an asset never overlap and each starts after the last repair ended; the run in force is rebuilt for every reading from the run table, and every label (run age, life used, remaining life, true remaining life, fails within 30 days) recomputes from it; no sensor value while the asset is down; pressure only on pumps and compressors, oil debris only on gearboxes and compressors; ISO 10816-3 zone, alarm and alarm reason recompute from the readings; one work order per run ended in the window, its type from how the run ended, its times from the run, its costs re-adding; inspection findings recompute from the measured vibration and follow each asset's PM interval; the physics above: P-F curve, mode-specific signals, load and climate against life, criticality against catches, sudden failures, false alarm rate.
- Can I use it commercially?
- Yes: in any project, course, benchmark, demo or product, commercial or not. You may not resell or redistribute the dataset itself as a dataset.
- What should I know before using it?
- ISO 10816-3 zone limits (2.3 / 4.5 / 7.1 mm/s) are applied to every asset for simplicity; real limits depend on machine class and foundation. true_rul_days and true_failure_at are ground truth no plant could know; use rul_days for a realistic label and the true values to study how censoring biases a model. pf_onset is the fraction of the true life at which degradation became visible (about 0.985 for a sudden failure). Lives are compressed so two years hold enough failures to learn from: the median life of a failure mode is 260 to 760 days, several times shorter than most plants see (a well-kept bearing or gear set often runs for years). The shape of each life (the P-F curve, the load and climate effects, the censoring) is what the data is built to teach, not its length; do not quote the MTBF as a benchmark. lost_production_cost is downtime hours times the site's hourly rate, scaled by criticality (A 1, B 0.45, C 0.12). Sites, assets and models are invented; no real plant, manufacturer or equipment record is represented.
- Can I get a bigger or different version?
- Yes. RECIPE.json runs in Misata Studio or through Misata's MCP server to make a variant, or ask us to build one to your spec.
Want it bigger, in another setting, or with your own columns? Have us build it or make it in Studio.
More premium datasets
All 4- readings, rows per monthJan 2025 – Jun 2025
3 plants, 36 CNC machines, six months
Six months of control charts with the ground truth underneath
- Rows
- 504,786
- Tables
- 15
- Checks
- 78 of 78
- fact_store_item_day, rows per monthFeb 2024 – Jan 2026
Grocery chain, 10 stores, fiscal 2024 and 2025
Two fiscal years of a grocery chain, from the shelf to the receipt
- Rows
- 2,780,138
- Tables
- 14
- Checks
- 110 of 110
- claims, rows per monthJan 2025 – Dec 2025
10 payers, 8 facilities, calendar 2025
A year of revenue cycle with the reason behind every denial
- Rows
- 356,609
- Tables
- 8
- Checks
- 87 of 87

