One prompt, two ways to make synthetic data
Tonic Fabricate and Claude + Misata on the same brief: a German coffee chain, a year of data, 15 realism checks.
Run on 1 October 2026, published 2 October 2026
This is not a claim that Misata is better than Fabricate.
It is a test of two different ways of generating synthetic data, run once each on one prompt. Fabricate's agent plans the data and writes a generator for each table. In the other approach, Claude writes one program describing the whole business, and Misata's engine runs it.
The short answer
Fabricate did the more complete job. It asked 8 sensible questions, covered every part of the brief, planned and listed 14 deliberate errors, and produced a strong understaffing effect. Its weak spots are in the detail: every item sells about equally, the menu does not change with the hour, rosters overlap, and comments repeat. The world program got most of those details right, because every table comes out of one simulated business. It forgot the deliberate errors, and its comments are templated. Costs were in the same range: $4.45 of credits for Fabricate, and about $3.68 of model time for Claude.
Tonic Fabricate
6 pass 5 weak 4 fail
- About 14 minutes, 8 questions answered
- $4.45 of credits
- 14 deliberate errors, as asked
Claude + Misata world program
13 pass 1 weak 1 fail
- 12 minutes writing, 75 s to generate
- About $3.68 of model time
- No deliberate errors (missed)
Scores are from our own 15 checks, out of 15. One run each.
The two approaches
Fabricate: an agent writes a generator for each table
You chat with Fabricate's agent. It asks questions, shows a plan that includes the deliberate errors it will add, then writes a generator for each table and runs it into a hosted database. It checks its own work with SQL, and you export the result. The model does the planning and writes the code; the code makes the rows.
Claude + Misata: one program for the whole business
Claude writes one Misata blueprint describing the whole business: shops and their hours, staff and contracts, who works when, how many customers arrive each hour, what they buy at that hour and season, and how long they wait given who is on duty. Misata's engine runs it with no model calls, and every table comes out of the same simulated days.
So both use a model to write a program and let the program make the rows. The difference is the shape of the program: one generator per table, or one model of the business that every table is read from. The second approach is not Misata Studio's one-click flow today. It was written in a coding session to test the idea.
The prompt, word for word
Build a complete operational database for a small invented German coffee chain with four shops in different German cities, covering about a year, in EUR with timestamps in Europe/Berlin. Include the shops, staff and their shifts, the menu, customer transactions with line items, and customer feedback and staff incidents. Transactions peak in a morning and an evening commuter rush and each shop has its own daily rhythm. Shops keep fixed opening hours. Two of the shops are understaffed at peak times, which should show up in slower service, worse customer feedback and more staff incidents. Add a handful of deliberate data errors.
This is the brief both sides were given.
How we scored
- pass The check holds, apart from planned errors.
- weak Mostly right, or right with a visible flaw a careful analyst would notice.
- fail Wrong in a way that would mislead anyone using the data.
The brief asked for deliberate errors, so anything a tool said it planned is credited, not counted against it. Peak hours are found from each shop's own data, not assumed. Dates are compared with dates and times with times. Every number below is measured on all the rows from each run. The one exception is the column-dependency check, which trains on a sample of up to 20,000 rows per table.
Cost and time
Prices are API list prices in October 2026. The Claude figure comes from the token counts on each call in the session log. A working session carries everything said before it, so most of the cost was re-reading that context. A short, dedicated call would carry far less, but we have not measured one.
The 15 checks
1. Everything the brief asked for
A dataset that skips part of the request has failed before realism even matters.
Tonic Fabricate pass
Every table asked for, plus a customers table. Each shop has its own rhythm and fixed hours, the understaffing shows up, and there are 14 deliberate errors, each one listed in its plan.
Claude + Misata fail
Every table asked for, plus a daily ledger for each shop. But no deliberate errors. Misata's blueprint can declare them, but this program never did, so it missed part of the brief.
2. A menu a café could print
Every sale points at the menu, so a broken menu spreads into every table that touches it.
Tonic Fabricate weak
33 rows of sensible German café items. A duplicate cappuccino and a negative-priced oat-milk surcharge are planned errors, and we credit them as such. But the negative surcharge was then sold on 18,039 lines, and 3,107 sales are made of nothing else, so their totals fall below zero.
Claude + Misata pass
35 distinct items, with the oat-milk surcharge charged on the line it belongs to. Two German VAT rates in use: 19% on 410,593 lines and 7% on 260,438.
3. Each sale's total equals its lines
The first check any analyst runs on a sales table.
Tonic Fabricate pass
100%, apart from 5 totals that are €2.50 off, which were planned.
Claude + Misata pass
100%.
4. A plausible basket
A coffee-shop sale is usually a drink, sometimes with something to eat.
Tonic Fabricate weak
2.46 items a sale, median €8.80, with merchandise on about 10% of lines (53,558). High for a coffee shop, though not impossible.
Claude + Misata pass
1.8 items a sale, median €6.30, and no sale at zero or below.
5. Best-sellers lead
In real sales a few items sell far more than the rest. A flat menu is a sign that items were picked at random.
Tonic Fabricate fail
Every regular item sold about 18,000 times. Cappuccino reached 36,018 only because it is on the menu twice.
Claude + Misata pass
Cappuccino leads with 84,981 lines, 6.8 times the median item. Seasonal specials, the soup of the day and lemonade sell least.
6. The menu changes with the hour
Breakfast pastries, lunch and afternoon cake are what make a café's day look real.
Tonic Fabricate fail
The share of each category is the same in every part of the day, to the percent. Soup was sold before 11:00 7,067 times.
Claude + Misata pass
Bakery is 22% of lines before 11:00 and 9% from 11:00 to 15:00. Lunch is 12% from 11:00 to 15:00, and cake 9% from 15:00. No soup before 11:00, and no butter pretzels after 13:00.
7. The menu changes with the season
Iced coffee in January and Christmas cake in June are giveaways.
Tonic Fabricate weak
Seasonal specials appear only in their season: no pumpkin spice in June. But cold coffee only moves from about 10% of lines to 13% in summer, and 4,228 iced drinks were sold in January.
Claude + Misata pass
Cold drinks are 3.5 to 4% of lines from October to March and 12.7 to 12.9% from May to August. Iced latte and cold brew are sold only from April to September, stollen only in November and December, and rhubarb cake from April to June.
8. Timestamps look recorded
A till records the second. Times all stored to the minute suggest the data was generated.
Tonic Fabricate weak
Every time on a sale, review, shift and incident falls on :00 seconds: stored to the minute.
Claude + Misata pass
1.7% of sales fall on :00 seconds, about what chance gives (1 in 60). Shift and incident times fall on the minute, as rosters do.
9. Shifts a person could work
An impossible roster breaks every staffing question you might ask of the data.
Tonic Fabricate fail
4,916 overlapping shifts, against 1 planned. 6,380 cases of one person working two shifts on the same day, and 3,453 person-days over 10 hours, the longest 19 hours. Shifts run from 1.5 to 14 hours.
Claude + Misata pass
Shifts of 5 to 8 hours, no overlaps, and nobody works twice in a day. Shifts per year follow the contract: median 272 for full-time, 156 for part-time and 73 for mini-jobs.
10. The server was employed and on shift
A sale served by someone who was not working that day cannot have happened.
Tonic Fabricate pass
99.42% of sales were served by someone on shift at the time, none after the server's leaving date. 6 sales have no server, which was planned.
Claude + Misata pass
99.38% of sales were served by someone on shift at the time, and none by anyone outside their employment.
11. Understaffing works as a cause
The heart of the brief. One cause should show up in three different tables: waits, reviews and incidents.
Tonic Fabricate pass
At each shop's own four busiest hours, the two understaffed shops have median waits of 175 and 171 seconds against 104 and 105; off-peak waits are 88 to 94 seconds everywhere. 14.1 and 16.5 incidents per 100 shifts against 4.2 and 4.3. Average rating 3.62 against 4.44. A strong, clean effect.
Claude + Misata pass
At each shop's own four busiest hours, the understaffed shops wait 194 and 128 seconds against 104 and 95; off-peak waits are 92 to 105 seconds. 3.0 and 6.2 incidents per 100 shifts against 2.0 and 2.3. Average ratings 4.11 and 4.34 against 4.45 and 4.43. Real, but milder than Fabricate's, and weak at one of the two shops.
12. Ratings look like reviews
Real ratings pile up at 4 and 5 with a small bump at 1, and waiting explains only part of a rating.
Tonic Fabricate weak
The one rating of 6 is a planned error. The rest climb steadily from 1 to 5 (908, 1,024, 2,083, 5,078 and 8,166), with no bump at 1. Rating follows wait almost exactly (correlation −0.90).
Claude + Misata pass
The real-review shape: 230, 266, 137, 2,272 and 3,281 for 1 to 5 stars. Rating against wait is −0.27, so waiting matters but is not the only thing.
13. Comments vary
Repeated text is the quickest giveaway in any synthetic dataset.
Tonic Fabricate fail
16 distinct comments across 11,401.
Claude + Misata weak
439 distinct comments across 2,114, assembled from templates. Better, but still visibly templated.
14. Columns depend on each other
In real data, columns move together. We train a model to tell each table from a copy with every column shuffled on its own: 0.5 means it cannot (the columns are independent), and close to 1.0 means real structure.
Tonic Fabricate pass
Sale lines 0.988, sales 0.999, shifts 1.0, feedback 0.766, incidents 0.818 (customers 0.461).
Claude + Misata pass
Sale lines 1.0, sales 0.989, shifts 1.0, shop days 0.991, feedback 0.742 (incidents 0.496, on only 212 rows).
15. Each shop is its own place
Four shops in four cities should not share one set of hours and one rhythm.
Tonic Fabricate pass
Each shop has its own hours and rhythm. A Hamburg shop with a Berlin postcode is a planned error. One shop closes at weekends, which is unusual for a café but possible.
Claude + Misata pass
Each shop has its own weekday and weekend hours and its own daily rhythm. The shop by a main station peaks at the morning and evening commute; the one in an office district peaks in the morning and at lunch.
Tally: Fabricate 6 pass, 5 weak, 4 fail. Claude + Misata 13 pass, 1 weak, 1 fail.
What each approach does well
Fabricate
- Asks before it builds, and its questions were good ones.
- Covers the whole brief, and plans and lists its deliberate errors up front.
- A clean, strong cause and effect across waits, ratings and incidents.
- A finished product: hosted database, self-checks and export, with no code to read.
Where it slipped: the details. Items sell equally at every hour, cold drinks barely follow the season, rosters overlap, times stop at the minute, and comments repeat.
Claude + Misata world program
- One model of the business, so sales, shifts, waits and reviews agree by construction.
- Rhythm by hour, weekday and season, with best-sellers and review shapes that look like real data.
- Seeded: the same program and seed give identical data, and new copies cost no model time.
- Fast to run once written: a full year in 75 seconds on a laptop.
Where it slipped: no deliberate errors, templated comments, and a milder understaffing effect at one shop.
What we learned
A model does not need to write rows, or even one generator per table. It needs to describe the world well once: who is there, when they work, what customers do at each hour and season, and what causes what. A deterministic engine can then make every table from that one description. The details that give synthetic data away (flat best-sellers, an identical menu at 8:00 and 16:00, impossible rosters) all disappeared in this run once the tables came from one simulated business. The things the world program missed, deliberate errors and varied comments, are the things that description left out.

