One prompt, two ways to make synthetic data

Tonic Fabricate and Claude + Misata on the same brief: a German coffee chain, a year of data, 15 realism checks.

Run on 1 October 2026, published 2 October 2026

This is not a claim that Misata is better than Fabricate.

It is a test of two different ways of generating synthetic data, run once each on one prompt. Fabricate's agent plans the data and writes a generator for each table. In the other approach, Claude writes one program describing the whole business, and Misata's engine runs it.

The short answer

Fabricate did the more complete job. It asked 8 sensible questions, covered every part of the brief, planned and listed 14 deliberate errors, and produced a strong understaffing effect. Its weak spots are in the detail: every item sells about equally, the menu does not change with the hour, rosters overlap, and comments repeat. The world program got most of those details right, because every table comes out of one simulated business. It forgot the deliberate errors, and its comments are templated. Costs were in the same range: $4.45 of credits for Fabricate, and about $3.68 of model time for Claude.

Tonic Fabricate

6 pass 5 weak 4 fail

  • About 14 minutes, 8 questions answered
  • $4.45 of credits
  • 14 deliberate errors, as asked

Claude + Misata world program

13 pass 1 weak 1 fail

  • 12 minutes writing, 75 s to generate
  • About $3.68 of model time
  • No deliberate errors (missed)

Scores are from our own 15 checks, out of 15. One run each.

The two approaches

Fabricate: an agent writes a generator for each table

You chat with Fabricate's agent. It asks questions, shows a plan that includes the deliberate errors it will add, then writes a generator for each table and runs it into a hosted database. It checks its own work with SQL, and you export the result. The model does the planning and writes the code; the code makes the rows.

Claude + Misata: one program for the whole business

Claude writes one Misata blueprint describing the whole business: shops and their hours, staff and contracts, who works when, how many customers arrive each hour, what they buy at that hour and season, and how long they wait given who is on duty. Misata's engine runs it with no model calls, and every table comes out of the same simulated days.

So both use a model to write a program and let the program make the rows. The difference is the shape of the program: one generator per table, or one model of the business that every table is read from. The second approach is not Misata Studio's one-click flow today. It was written in a coding session to test the idea.

The prompt, word for word

Build a complete operational database for a small invented German coffee chain with four shops in different German cities, covering about a year, in EUR with timestamps in Europe/Berlin. Include the shops, staff and their shifts, the menu, customer transactions with line items, and customer feedback and staff incidents. Transactions peak in a morning and an evening commuter rush and each shop has its own daily rhythm. Shops keep fixed opening hours. Two of the shops are understaffed at peak times, which should show up in slower service, worse customer feedback and more staff incidents. Add a handful of deliberate data errors.

This is the brief both sides were given.

How we scored

  • pass The check holds, apart from planned errors.
  • weak Mostly right, or right with a visible flaw a careful analyst would notice.
  • fail Wrong in a way that would mislead anyone using the data.

The brief asked for deliberate errors, so anything a tool said it planned is credited, not counted against it. Peak hours are found from each shop's own data, not assumed. Dates are compared with dates and times with times. Every number below is measured on all the rows from each run. The one exception is the column-dependency check, which trains on a sample of up to 20,000 rows per table.

Cost and time

Who did the workFabricate: Fabricate's agent, v4.32.0, on the free trial. The agent model setting showed Sonnet 5, High; the validation agent was off.Claude + Misata: Claude, working in Claude Code. It wrote one Misata blueprint for the whole business, about 360 lines, and ran it on Misata's engine.
Questions askedFabricate: 8, in two rounds, answered by a personClaude + Misata: None
TimeFabricate: About 14 minutes from prompt to export (17:35 to 17:49 UTC on 1 October 2026), including the time a person took to answerClaude + Misata: 12 minutes of writing and fixing (18:16 to 18:28 UTC), over 57 model calls. Before that, 7 minutes of reading, mostly about Misata's existing pipeline.
Making the rowsFabricate: Part of the agent's run, into a hosted databaseClaude + Misata: 75 seconds on a laptop, with no model calls
Model costFabricate: $4.45 of credits, out of the $5 free allowanceClaude + Misata: About $3.68 at API list prices for the 12 minutes, or about $6.01 counting the 7 minutes before
What that cost isFabricate: A price to the customer, so it includes Tonic's marginClaude + Misata: Raw token cost: about 60,000 tokens of output, 6.6 million tokens of earlier context re-read from cache across the 57 calls, and 144,000 tokens written to cache. Most of it is re-reading a long session.
Another copy of the dataFabricate: Not measuredClaude + Misata: 75 seconds and no model cost. The same seed gives identical data.
Rows producedFabricate: 269,344 sales, 551,836 sale lines, 9,833 shifts, 17,260 reviews, 690 incidentsClaude + Misata: 373,432 sales, 671,031 sale lines, 6,824 shifts, 6,186 reviews, 212 incidents

Prices are API list prices in October 2026. The Claude figure comes from the token counts on each call in the session log. A working session carries everything said before it, so most of the cost was re-reading that context. A short, dedicated call would carry far less, but we have not measured one.

The 15 checks

  1. 1. Everything the brief asked for

    A dataset that skips part of the request has failed before realism even matters.

    Tonic Fabricate pass

    Every table asked for, plus a customers table. Each shop has its own rhythm and fixed hours, the understaffing shows up, and there are 14 deliberate errors, each one listed in its plan.

    Claude + Misata fail

    Every table asked for, plus a daily ledger for each shop. But no deliberate errors. Misata's blueprint can declare them, but this program never did, so it missed part of the brief.

  2. 2. A menu a café could print

    Every sale points at the menu, so a broken menu spreads into every table that touches it.

    Tonic Fabricate weak

    33 rows of sensible German café items. A duplicate cappuccino and a negative-priced oat-milk surcharge are planned errors, and we credit them as such. But the negative surcharge was then sold on 18,039 lines, and 3,107 sales are made of nothing else, so their totals fall below zero.

    Claude + Misata pass

    35 distinct items, with the oat-milk surcharge charged on the line it belongs to. Two German VAT rates in use: 19% on 410,593 lines and 7% on 260,438.

  3. 3. Each sale's total equals its lines

    The first check any analyst runs on a sales table.

    Tonic Fabricate pass

    100%, apart from 5 totals that are €2.50 off, which were planned.

    Claude + Misata pass

    100%.

  4. 4. A plausible basket

    A coffee-shop sale is usually a drink, sometimes with something to eat.

    Tonic Fabricate weak

    2.46 items a sale, median €8.80, with merchandise on about 10% of lines (53,558). High for a coffee shop, though not impossible.

    Claude + Misata pass

    1.8 items a sale, median €6.30, and no sale at zero or below.

  5. 5. Best-sellers lead

    In real sales a few items sell far more than the rest. A flat menu is a sign that items were picked at random.

    Tonic Fabricate fail

    Every regular item sold about 18,000 times. Cappuccino reached 36,018 only because it is on the menu twice.

    Claude + Misata pass

    Cappuccino leads with 84,981 lines, 6.8 times the median item. Seasonal specials, the soup of the day and lemonade sell least.

  6. 6. The menu changes with the hour

    Breakfast pastries, lunch and afternoon cake are what make a café's day look real.

    Tonic Fabricate fail

    The share of each category is the same in every part of the day, to the percent. Soup was sold before 11:00 7,067 times.

    Claude + Misata pass

    Bakery is 22% of lines before 11:00 and 9% from 11:00 to 15:00. Lunch is 12% from 11:00 to 15:00, and cake 9% from 15:00. No soup before 11:00, and no butter pretzels after 13:00.

  7. 7. The menu changes with the season

    Iced coffee in January and Christmas cake in June are giveaways.

    Tonic Fabricate weak

    Seasonal specials appear only in their season: no pumpkin spice in June. But cold coffee only moves from about 10% of lines to 13% in summer, and 4,228 iced drinks were sold in January.

    Claude + Misata pass

    Cold drinks are 3.5 to 4% of lines from October to March and 12.7 to 12.9% from May to August. Iced latte and cold brew are sold only from April to September, stollen only in November and December, and rhubarb cake from April to June.

  8. 8. Timestamps look recorded

    A till records the second. Times all stored to the minute suggest the data was generated.

    Tonic Fabricate weak

    Every time on a sale, review, shift and incident falls on :00 seconds: stored to the minute.

    Claude + Misata pass

    1.7% of sales fall on :00 seconds, about what chance gives (1 in 60). Shift and incident times fall on the minute, as rosters do.

  9. 9. Shifts a person could work

    An impossible roster breaks every staffing question you might ask of the data.

    Tonic Fabricate fail

    4,916 overlapping shifts, against 1 planned. 6,380 cases of one person working two shifts on the same day, and 3,453 person-days over 10 hours, the longest 19 hours. Shifts run from 1.5 to 14 hours.

    Claude + Misata pass

    Shifts of 5 to 8 hours, no overlaps, and nobody works twice in a day. Shifts per year follow the contract: median 272 for full-time, 156 for part-time and 73 for mini-jobs.

  10. 10. The server was employed and on shift

    A sale served by someone who was not working that day cannot have happened.

    Tonic Fabricate pass

    99.42% of sales were served by someone on shift at the time, none after the server's leaving date. 6 sales have no server, which was planned.

    Claude + Misata pass

    99.38% of sales were served by someone on shift at the time, and none by anyone outside their employment.

  11. 11. Understaffing works as a cause

    The heart of the brief. One cause should show up in three different tables: waits, reviews and incidents.

    Tonic Fabricate pass

    At each shop's own four busiest hours, the two understaffed shops have median waits of 175 and 171 seconds against 104 and 105; off-peak waits are 88 to 94 seconds everywhere. 14.1 and 16.5 incidents per 100 shifts against 4.2 and 4.3. Average rating 3.62 against 4.44. A strong, clean effect.

    Claude + Misata pass

    At each shop's own four busiest hours, the understaffed shops wait 194 and 128 seconds against 104 and 95; off-peak waits are 92 to 105 seconds. 3.0 and 6.2 incidents per 100 shifts against 2.0 and 2.3. Average ratings 4.11 and 4.34 against 4.45 and 4.43. Real, but milder than Fabricate's, and weak at one of the two shops.

  12. 12. Ratings look like reviews

    Real ratings pile up at 4 and 5 with a small bump at 1, and waiting explains only part of a rating.

    Tonic Fabricate weak

    The one rating of 6 is a planned error. The rest climb steadily from 1 to 5 (908, 1,024, 2,083, 5,078 and 8,166), with no bump at 1. Rating follows wait almost exactly (correlation −0.90).

    Claude + Misata pass

    The real-review shape: 230, 266, 137, 2,272 and 3,281 for 1 to 5 stars. Rating against wait is −0.27, so waiting matters but is not the only thing.

  13. 13. Comments vary

    Repeated text is the quickest giveaway in any synthetic dataset.

    Tonic Fabricate fail

    16 distinct comments across 11,401.

    Claude + Misata weak

    439 distinct comments across 2,114, assembled from templates. Better, but still visibly templated.

  14. 14. Columns depend on each other

    In real data, columns move together. We train a model to tell each table from a copy with every column shuffled on its own: 0.5 means it cannot (the columns are independent), and close to 1.0 means real structure.

    Tonic Fabricate pass

    Sale lines 0.988, sales 0.999, shifts 1.0, feedback 0.766, incidents 0.818 (customers 0.461).

    Claude + Misata pass

    Sale lines 1.0, sales 0.989, shifts 1.0, shop days 0.991, feedback 0.742 (incidents 0.496, on only 212 rows).

  15. 15. Each shop is its own place

    Four shops in four cities should not share one set of hours and one rhythm.

    Tonic Fabricate pass

    Each shop has its own hours and rhythm. A Hamburg shop with a Berlin postcode is a planned error. One shop closes at weekends, which is unusual for a café but possible.

    Claude + Misata pass

    Each shop has its own weekday and weekend hours and its own daily rhythm. The shop by a main station peaks at the morning and evening commute; the one in an office district peaks in the morning and at lunch.

Tally: Fabricate 6 pass, 5 weak, 4 fail. Claude + Misata 13 pass, 1 weak, 1 fail.

What each approach does well

Fabricate

  • Asks before it builds, and its questions were good ones.
  • Covers the whole brief, and plans and lists its deliberate errors up front.
  • A clean, strong cause and effect across waits, ratings and incidents.
  • A finished product: hosted database, self-checks and export, with no code to read.

Where it slipped: the details. Items sell equally at every hour, cold drinks barely follow the season, rosters overlap, times stop at the minute, and comments repeat.

Claude + Misata world program

  • One model of the business, so sales, shifts, waits and reviews agree by construction.
  • Rhythm by hour, weekday and season, with best-sellers and review shapes that look like real data.
  • Seeded: the same program and seed give identical data, and new copies cost no model time.
  • Fast to run once written: a full year in 75 seconds on a laptop.

Where it slipped: no deliberate errors, templated comments, and a milder understaffing effect at one shop.

What we learned

A model does not need to write rows, or even one generator per table. It needs to describe the world well once: who is there, when they work, what customers do at each hour and season, and what causes what. A deterministic engine can then make every table from that one description. The details that give synthetic data away (flat best-sellers, an identical menu at 8:00 and 16:00, impossible rosters) all disappeared in this run once the tables came from one simulated business. The things the world program missed, deliberate errors and varied comments, are the things that description left out.

Questions

Is Misata better than Tonic Fabricate?
This page does not say so, and that is not what it tests. It runs one prompt through two different ways of making synthetic data and reports what came out. Fabricate is a complete, polished product that asks good questions and covered the whole brief. The world program is an experiment written in a coding session, and Misata Studio does not do it with one click today.
Neither approach uses the model to write every row, so why do the costs look so similar?
Both use a model to write a program, and the program makes the rows. Fabricate's agent writes a generator for each table, and the Claude session wrote one program for the whole business. Most of the Claude cost was re-reading a long working session's context on each of 57 calls, not writing the program itself. The two figures are also in different units: Fabricate's is a price to customers, ours is raw API list price.
What is a world program?
One declarative description of the whole business: the shops and their hours, who works which shift, how many customers arrive each hour, what they buy at that hour and season, and how long they wait given who is on duty. Misata's engine runs it with no model calls and makes every table from the same simulated events, so the tables agree with each other by construction. Understaffing is a cause in the program, not a label added afterwards.
Why did the world program fail the first check?
The brief asked for a handful of deliberate data errors. Misata's blueprint can declare them, but the program written for this test did not, so it missed part of the request. Fabricate planned 14 and listed each one.
Can I get the world program's result in Misata Studio today?
Not from a single prompt. Studio's current one-click flow is a different design, and it does not reach this level on this brief yet. A model writing the whole world once, and the engine running it, is the direction we are now working on.
Why are Fabricate's deliberate errors not counted against it?
The brief asked for them, and Fabricate listed each one in its plan. We credit planned errors and score only what was not planned. Where a planned error spread further than planned, as the negative surcharge did into 3,107 sales, we note it.
Can I run this test myself?
Yes. The prompt is on this page word for word, and every check is described with its method. Run the prompt in any tool and apply the same checks. Find peak hours from the data rather than assuming them, compare dates with dates, and leave out any errors the tool says it planned.