Premium · 240 machines at 4 plants, two years

Predictive Maintenance Fleet

Two years of sensor readings, failures, repairs and costs for 240 machines

179,029 rows in 6 linked tables: 240 rotating assets at 4 plants, 175,440 daily sensor readings, 771 asset lives (runs), 532 work orders and 2,042 route inspections, 1 January 2024 to 31 December 2025.

Why it exists

Real run-to-failure data is scarce: plants replace parts before they fail, failures are rare, and the few public sets are one machine on a test rig. This fleet has the whole picture a reliability team works with: many machine types and failure modes, sensors that respond to the mode that is developing, repairs that cut lives short (censoring), and the money. Every reading carries the true remaining life, even where the part was replaced before it could fail, so a remaining-life model can be scored against what would really have happened.

Uses

What it is built for

  • Remaining useful life (RUL) regression

    with rul_days or true_rul_days, and failure within 30 days classification with fails_within_30d (5.6% positive, the realistic imbalance).

  • Survival analysis

    239 runs are right-censored at the cut-off and 204 were ended early by a repair or replacement; Weibull, Cox and Kaplan-Meier on runs.csv with load, climate and failure mode as covariates.

  • Anomaly detection and alarm tuning

    the alarms are rule-based (ISO zone, temperature, pressure, current, debris); beat them with a model, and measure false alarms against lead time.

  • Failure mode diagnosis

    from the sensor signature: vibration for bearings, imbalance and misalignment, temperature for lubrication, pressure for impellers and seals, current for windings, oil debris for gears.

  • Maintenance strategy and cost

    run-to-failure against condition-based against age-based replacement; breakdowns cost a median $30,991 against $9,796 for a condition-based repair.

  • Dashboards and teaching

    MTBF, MTTR, availability, backlog, PM compliance, cost by site and criticality in Power BI, Tableau or SQL; a reliability engineering course or a CMMS demo.

In the data

What the full files show

Computed from every row of the full dataset when it was packaged, not drawn as a target. The free preview is a slice of the same data.

Vibration through a life that ends in failure

0.04.08.00%15%30%45%60%75%90%
  • Gave warning
  • Sudden
Median vibration RMS (mm/s) by the share of the run's true life used, for runs that ended in a failure. Failures with warning signs rise along a P-F curve near the end; sudden ones stay flat until they break.

Chance of failing within 30 days, by ISO 10816-3 zone

  • Zone A2.3%
  • Zone B8.3%
  • Zone C26%
  • Zone D85.3%
Share of daily readings, while running, whose machine truly fails within 30 days (fails_within_30d), grouped by the vibration zone the reading falls in. Zone A is newly commissioned condition, D is damage.

Failures by mode

  • bearing wear144
  • shaft misalignment81
  • seal leak72
  • rotor imbalance57
  • lubrication breakdown47
  • impeller erosion36
  • gear tooth wear34
  • winding insulation28
Runs that ended in a failure, by the true failure mode in runs. Each mode moves its own signals: bearings and lubrication heat up, impellers and seals lose pressure.

What a work order costs, by type

  • corrective (breakdown)$30,991
  • corrective (planned, condition-based)$9,796
  • preventive replacement$8,348
Median total cost (parts, labour and lost production) per work order, from work_orders. A breakdown keeps the machine down far longer than planned work, and lost production is most of the bill.

Answer key

The truth, in its own columns

What a model, a control chart or an analyst is trying to find ships next to the data, so you can score an answer instead of guessing at it. Each line is measured from the full files.

  • runs.failure_modeThe true failure mode of the life, which sets the signals that move.8 values: bearing wear, shaft misalignment, seal leak…
  • runs.true_life_daysHow long the life lasts to failure if nothing intervenes; a run caught or replaced early ends before it.52.8 to 3,979.8, mean 428.9
  • runs.true_failure_atWhen the machine would fail, including for runs that were caught or renewed first.2024-01-11 19:12:00 to 2035-08-01 22:00:00
  • runs.warning_signsWhether the failure gives signs a sensor can see, or comes suddenly.true on 82.2% of rows
  • runs.pf_onsetThe share of the life at which the signals start to rise (the P on the P-F curve); near 1 for sudden failures.0.55 to 0.98, mean 0.69
  • readings.life_usedThe share of the run's true life used at the reading.0 to 1, mean 0.46, 0.78% empty
  • readings.rul_daysDays left until the run ends, however it ends.0 to 3,979.4, mean 289.6, 0.78% empty
  • readings.true_rul_daysDays left until the true failure: the remaining-useful-life target.0 to 3,979.4, mean 325.6, 0.78% empty
  • readings.fails_within_30dWhether the machine truly fails within 30 days: the classification target.true on 5.6% of rows

How it behaves

Measured on the files you download

Each of these is computed from the delivered rows when the dataset is packaged, not written as a target.

  • Signals stay flat for most of a life, then rise along a P-F curve: median vibration on vibration-type failures goes 1.62 -> 1.63 -> 2.50 -> 8.62 mm/s from new to the last 3% of life.
  • 91% of warned and 7% of sudden failures gave a sustained alarm first, median 46 days ahead. 25% of failures were sudden: some faults give no warning a sensor can see.
  • Harder-worked machines live shorter (bearings with the cube of load): Spearman -0.57 between load factor and true life; hot sites wear machines out sooner (median true life 275 against 333 days).
  • Near the end of a life, bearing temperature is 62 C against 42 C for bearing and lubrication failures against other modes; discharge pressure falls to 78% of nominal against 96% for impeller and seal failures.
  • Breakdowns keep a machine down a median 51 h against 6 h for planned work; lost production is $20,565,604 of the $22,049,444 maintenance cost.
  • Runs: 328 failures, 123 caught by monitoring, 81 age-based replacements, 239 still running at the cut-off. Sensor dropouts: about 0.8% of readings.

Audit

62 of 62 checks pass

Re-run on the delivered files by an independent script with plain pandas. The results ship in INTEGRITY.json.

  • every key resolves; one reading per asset per day for all 731 days; runs of an asset never overlap and each starts after the last repair ended;
  • the run in force is rebuilt for every reading from the run table, and every label (run age, life used, remaining life, true remaining life, fails within 30 days) recomputes from it;
  • no sensor value while the asset is down; pressure only on pumps and compressors, oil debris only on gearboxes and compressors; ISO 10816-3 zone, alarm and alarm reason recompute from the readings;
  • one work order per run ended in the window, its type from how the run ended, its times from the run, its costs re-adding; inspection findings recompute from the measured vibration and follow each asset's PM interval;
  • the physics above: P-F curve, mode-specific signals, load and climate against life, criticality against catches, sudden failures, false alarm rate.
All 62 checks
  • ✓ sites.site_id unique
  • ✓ assets.asset_id unique
  • ✓ runs.run_id unique
  • ✓ readings.reading_id unique
  • ✓ work_orders.work_order_id unique
  • ✓ inspections.inspection_id unique
  • ✓ assets.site_id -> sites: 0 orphans of 240
  • ✓ runs.asset_id -> assets: 0 orphans of 771
  • ✓ runs.site_id -> sites: 0 orphans of 771
  • ✓ readings.asset_id -> assets: 0 orphans of 175,440
  • ✓ readings.site_id -> sites: 0 orphans of 175,440
  • ✓ readings.run_id -> runs: 0 orphans of 174,063
  • ✓ work_orders.run_id -> runs: 0 orphans of 532
  • ✓ work_orders.asset_id -> assets: 0 orphans of 532
  • ✓ inspections.asset_id -> assets: 0 orphans of 2,042
  • ✓ a run's site and a reading's site are the asset's
  • ✓ one reading per asset per day, every day of both years: 731 days x 240 assets
  • ✓ a run's end = start + life; true failure = start + true life; back in service = end + downtime
  • ✓ a failure ends at the true life; a renewal before it
  • ✓ each run starts the moment the last repair ends
  • ✓ run numbers count up from 1 per asset
  • ✓ each asset is part-way through a run on the first day
  • ✓ no run starts after the cut-off
  • ✓ outcome at cut-off: running when the run outlives the data
  • ✓ every failure mode can happen to its kind of asset
  • ✓ run_id is the run in force at the reading, empty while the asset is down
  • ✓ running exactly when a run is in force; no run hours while down
  • ✓ rul_days = run end - reading
  • ✓ true_rul_days = true failure - reading (past a renewal too)
  • ✓ life_used = age / true life
  • ✓ fails_within_30d only in a run's last 30 days before a real failure
  • ✓ no sensor value while the asset is down
  • ✓ sensor dropouts are rare (about 1%): vibration_rms_mm_s 0.84%, bearing_temp_c 0.78%, motor_current_a 0.80%
  • ✓ pressure only on pumps and compressors, oil debris only on gearboxes and compressors
  • ✓ run hours are the asset's duty hours while running
  • ✓ ISO 10816-3 zone recomputes from vibration
  • ✓ alarm reason recomputes (vibration zone C/D, bearing over 85 C, pressure under 88% or current over 108% of expected at that load, oil debris over 1500/ml)
  • ✓ daily load wanders around each asset's own load factor: median gap 0.5 points
  • ✓ readings stay physical
  • ✓ one work order for every run that ended inside the window, none for others
  • ✓ work order type follows how the run ended
  • ✓ work starts when the run ends and closes when the asset is back
  • ✓ a breakdown is raised when it happens, planned work before
  • ✓ failure mode and component replaced are the run's
  • ✓ total cost = parts + labour + lost production
  • ✓ labour cost = hours x rate (breakdown overtime 135, planned 95)
  • ✓ found condition: failed on breakdowns, a confirmed defect on monitored catches
  • ✓ an inspection finds the asset down exactly when no run is in force
  • ✓ inspection finding recomputes from the measured vibration
  • ✓ route frequency follows each asset's PM interval: 6-12 per asset
  • ✓ vibration flat for most of the life, then rising toward a vibration failure: 1.62 -> 1.63 -> 2.50 -> 8.62
  • ✓ bearings and lubrication run hot near the end; other modes do not: 62 C against 42 C
  • ✓ impeller erosion and seal leaks lose pressure: 78% of nominal against 96%
  • ✓ hot sites wear equipment out sooner: median true life 275 against 333 days
  • ✓ harder-worked assets live shorter: Spearman -0.57
  • ✓ critical assets have more failures caught in time: A 49%, B 19%, C 8%
  • ✓ breakdowns keep an asset down far longer than planned work: 51 h against 6 h
  • ✓ failures with warning signs alarm first, weeks ahead (the P-F interval); sudden ones mostly do not: 91% of warned and 7% of sudden failures gave a sustained alarm first, median 46 days ahead; 25% of failures sudden
  • ✓ only a run with warning signs is caught by monitoring
  • ✓ false alarms happen on healthy machines, rarely: 0.17% of readings in the first half of a life
  • ✓ zone D is rare: 1.9% of readings
  • ✓ every asset ends the data running or under repair; runs are right-censored at the cut-off: 239 censored runs
  • ✓ runs: 771; failure 328, running at cut-off 239, caught by monitoring 123, age-based replacement 81
  • ✓ failure modes: bearing wear 32%, shaft misalignment 15%, seal leak 12%, lubrication breakdown 11%, rotor imbalance 11%, impeller erosion 7%, gear tooth wear 6%, winding insulation 6%
  • ✓ maintenance spend: $22,049,444, of it lost production $20,565,604

Tables

6 tables, 179,029 rows

Every table in the zip with what it holds, its rows and its columns. The bars are on a log scale, so the small reference tables still show.

  • readingsOne reading per asset per day for two years: vibration RMS and crest factor, bearing temperature, motor current, discharge pressure, oil debris, load, ambient, ISO 10816-3 zone, alarm and its reason; plus labels (run age, life used, remaining life to the run's end and to the true failure, fails within 30 days)175,440 rows · 24 cols
  • inspectionsRoute-based vibration checks at each asset's PM interval, with the measured value and the finding2,042 rows · 7 cols
  • runsEvery life of every asset, from install or repair to the next failure or renewal: failure mode, true life, whether it gave warning signs, how it ended, downtime, and whether it was still running at the cut-off771 rows · 19 cols
  • work_ordersOne per run that ended in the window: breakdown, condition-based or preventive, priority, raised/started/completed, component replaced, found condition, labour hours, parts, labour and lost-production cost532 rows · 19 cols
  • assetsPumps, motors, fans, gearboxes and compressors: type, criticality A/B/C, rated power and current, nominal pressure, install year, load factor, duty hours, online monitoring or route-based, PM interval240 rows · 15 cols
  • sitesFour plants in different climates: ambient temperature and its seasonal swing, spares lead time, cost of an hour of downtime4 rows · 7 cols

Explore

Every table, profiled

Each column's type, spread, empties and most common values, measured from the full CSVs. Switch to the first rows to see the data as it sits in the file.

readings.csv

175,440 rows · 24 columns · 3 foreign keys

reading_idprimary key
unique on every row
175,440 distinctno empties
asset_idforeign key
points to assets.asset_id
240 distinctno empties
site_idforeign key
points to sites.site_id
4 distinctno empties
reading_datedate
Jan 2024Dec 2025
2024-01-01 to 2025-12-31no empties
reading_atdate
Jan 2024Dec 2025
2024-01-01 to 2025-12-31no empties
run_idforeign key
points to runs.run_id
771 distinct0.8% empty
runningyes / no
true · 99%false · 0.8%
no empties
run_age_daysnumber
1.71,274.7
mean 251.6median 174.90 to 1,575.60.8% empty
life_usednumber
0.00430.99
mean 0.46median 0.450 to 10.8% empty
rul_daysnumber
1.72,649.3
mean 289.6median 182.40 to 3,979.40.8% empty
true_rul_daysnumber
2.72,649.3
mean 325.6median 212.40 to 3,979.40.8% empty
fails_within_30dyes / no
true · 5.6%false · 94%
no empties
run_hoursnumber
024
mean 22.48median 240 to 24no empties
ambient_cnumber
-12.138.4
mean 13.59median 14-18.6 to 44no empties
load_pctnumber
26.3100
mean 67.83median 69.315 to 1000.8% empty
vibration_rms_mm_snumber
0.878.64
mean 2.05median 1.680.72 to 13.271.6% empty
vibration_crest_factornumber
2.86.03
mean 3.2median 3.082.8 to 6.491.6% empty
bearing_temp_cnumber
11.772.1
mean 40.17median 40.13.2 to 99.91.6% empty
motor_current_anumber
19.4479.1
mean 96.75median 75.314 to 566.61.6% empty
discharge_pressure_barnumber
3.839.75
mean 6.93median 6.962.96 to 9.9745% empty
oil_particles_per_mlnumber
1092,455
mean 270.9median 18766 to 4,64666% empty
iso_zonecategory
  • A
    79%
  • B
    15%
  • C
    3.8%
  • D
    1.9%
4 values1.6% empty
alarm_reasoncategory
  • vibration
    82%
  • pressure loss
    14%
  • oil debris
    2.2%
  • motor current
    2.1%
  • bearing temperature
    0.1%
5 values93% empty
alarmyes / no
true · 6.8%false · 93%
no empties

Keys

Every join resolves

10 foreign keys, 530,363 references checked against the table each one points at. None points at a row that does not exist.

assets

  • site_idsites.site_id0 orphans

inspections

  • asset_idassets.asset_id0 orphans

readings

  • asset_idassets.asset_id0 orphans
  • site_idsites.site_id0 orphans
  • run_idruns.run_id0 orphans

runs

  • asset_idassets.asset_id0 orphans
  • site_idsites.site_id0 orphans

work_orders

  • run_idruns.run_id0 orphans
  • asset_idassets.asset_id0 orphans
  • site_idsites.site_id0 orphans

In the zip

What you get

  • CSVs: one file per table, with a header row and ISO dates.
  • README.md: what each table holds, how the data behaves (measured), what was checked and what to know.
  • INTEGRITY.json: the 62 audit checks and their results.
  • RECIPE.json: the Misata blueprint that made these exact rows, seed 20240101.
All 9 files
  • sites.csv222 B
  • assets.csv20 KB
  • runs.csv145 KB
  • readings.csv23.2 MB
  • work_orders.csv109 KB
  • inspections.csv184 KB
  • README.md8 KB
  • INTEGRITY.json9 KB
  • RECIPE.json27 KB

Before you use it

Things to know

  • ISO 10816-3 zone limits (2.3 / 4.5 / 7.1 mm/s) are applied to every asset for simplicity; real limits depend on machine class and foundation.
  • true_rul_days and true_failure_at are ground truth no plant could know; use rul_days for a realistic label and the true values to study how censoring biases a model.
  • pf_onset is the fraction of the true life at which degradation became visible (about 0.985 for a sudden failure).
  • Lives are compressed so two years hold enough failures to learn from: the median life of a failure mode is 260 to 760 days, several times shorter than most plants see (a well-kept bearing or gear set often runs for years). The shape of each life (the P-F curve, the load and climate effects, the censoring) is what the data is built to teach, not its length; do not quote the MTBF as a benchmark.
  • lost_production_cost is downtime hours times the site's hourly rate, scaled by criticality (A 1, B 0.45, C 0.12).
  • Sites, assets and models are invented; no real plant, manufacturer or equipment record is represented.

Questions

Before you buy

Is this real data?
No. This is synthetic data generated by software. No row describes a real person, company, patient, store, machine or transaction. The patterns are modelled to be realistic and the statistics quoted are measured on these files, but they do not describe any real population or market. Use it for learning, testing, demos, benchmarks and prototyping, not as evidence about the real world. Provided as is, without warranty.
What do I get?
One zip of 6.1 MB: 6 linked tables and 179,029 rows as CSV, a README of what each table holds and how the data behaves, INTEGRITY.json with the 62 audit checks, and RECIPE.json, the Misata blueprint that made these exact rows (seed 20240101).
Can I try it before buying?
Yes. The free preview (303 KB) is a slice of the same data with the keys intact, plus the README, so you can load it and check it fits before you pay.
How was it checked?
62 of 62 checks pass, re-run on the delivered files by an independent script with plain pandas: every key resolves; one reading per asset per day for all 731 days; runs of an asset never overlap and each starts after the last repair ended; the run in force is rebuilt for every reading from the run table, and every label (run age, life used, remaining life, true remaining life, fails within 30 days) recomputes from it; no sensor value while the asset is down; pressure only on pumps and compressors, oil debris only on gearboxes and compressors; ISO 10816-3 zone, alarm and alarm reason recompute from the readings; one work order per run ended in the window, its type from how the run ended, its times from the run, its costs re-adding; inspection findings recompute from the measured vibration and follow each asset's PM interval; the physics above: P-F curve, mode-specific signals, load and climate against life, criticality against catches, sudden failures, false alarm rate.
Can I use it commercially?
Yes: in any project, course, benchmark, demo or product, commercial or not. You may not resell or redistribute the dataset itself as a dataset.
What should I know before using it?
ISO 10816-3 zone limits (2.3 / 4.5 / 7.1 mm/s) are applied to every asset for simplicity; real limits depend on machine class and foundation. true_rul_days and true_failure_at are ground truth no plant could know; use rul_days for a realistic label and the true values to study how censoring biases a model. pf_onset is the fraction of the true life at which degradation became visible (about 0.985 for a sudden failure). Lives are compressed so two years hold enough failures to learn from: the median life of a failure mode is 260 to 760 days, several times shorter than most plants see (a well-kept bearing or gear set often runs for years). The shape of each life (the P-F curve, the load and climate effects, the censoring) is what the data is built to teach, not its length; do not quote the MTBF as a benchmark. lost_production_cost is downtime hours times the site's hourly rate, scaled by criticality (A 1, B 0.45, C 0.12). Sites, assets and models are invented; no real plant, manufacturer or equipment record is represented.
Can I get a bigger or different version?
Yes. RECIPE.json runs in Misata Studio or through Misata's MCP server to make a variant, or ask us to build one to your spec.

Want it bigger, in another setting, or with your own columns? Have us build it or make it in Studio.

More premium datasets

All 4
  • readings, rows per monthJan 2025 – Jun 2025

    3 plants, 36 CNC machines, six months

    Six months of control charts with the ground truth underneath

    Rows
    504,786
    Tables
    15
    Checks
    78 of 78
    Profile and preview
  • fact_store_item_day, rows per monthFeb 2024 – Jan 2026

    Grocery chain, 10 stores, fiscal 2024 and 2025

    Two fiscal years of a grocery chain, from the shelf to the receipt

    Rows
    2,780,138
    Tables
    14
    Checks
    110 of 110
    Profile and preview
  • claims, rows per monthJan 2025 – Dec 2025

    10 payers, 8 facilities, calendar 2025

    A year of revenue cycle with the reason behind every denial

    Rows
    356,609
    Tables
    8
    Checks
    87 of 87
    Profile and preview