Free datasets

Sample data where the tables actually join.

Most sample data is one flat CSV, which is useless the moment you need a real join, a date dimension, or a number that has to add up. These are multi-table datasets with the foreign keys intact and the arithmetic reconciled. Free, no signup, public domain.

Every claim under each dataset is a check that was run against the actual files. Where something is imperfect, it says so.

2 datasets93,576 rows10 tablesCC0 public domain

Ecommerce storefront

A year of orders across five joined tables, with a real Q4 peak and totals that reconcile to the cent.

An online retailer's 2025: customers, a 300-SKU catalogue, orders, line items, and reviews. Every order total is exactly the sum of its own line items, so the joins and the arithmetic both hold when you check them. Demand rises into November and December because there are more orders, not because the orders got bigger, which is how real seasonality works.

TableRowsColumns
customers2,000customer_id, full_name, email, city, state, country, signup_date, segment112 signed up and never ordered
products300product_id, product_name, category, unit_price, unit_costevery SKU name distinct
orders11,081order_id, customer_id, order_date, status, channel, order_total
order_items31,000order_item_id, order_id, product_id, quantity, unit_price, line_total
reviews4,811review_id, product_id, customer_id, rating, review_text, review_date

What was verified

  • 0 orphaned foreign keys across all 5 relationships
  • 0 orders dated before their customer signed up
  • order_total equals the sum of its line items, exactly, for every order
  • 0 products priced at or below cost (margins run 22% to 64%)
  • Ratings are J-shaped (57% five-star, 7% one-star), not uniform
  • Top 10% of customers place 29.5% of orders, a realistic Pareto tail
  • 59% of prices end in .99, as real catalogues do
  • Every city belongs to its country (London and Newcastle appear under two, correctly)

Questions worth asking it

  • Which category carries the best margin, and is it the one selling most?
  • How much of revenue comes from the top 10% of customers?
  • Do low-rated products actually get returned more often?
  • What does the Q4 lift look like split by channel?

Regenerate it yourself with seed 20260721 for byte-identical output, or change the story and get a dataset nobody else has.

B2B SaaS subscription analytics

Accounts, seats, MRR, churn and support load, where company size actually drives the plan.

A B2B SaaS business with 1,200 customer accounts. Company size follows a power law, so most customers are small and a few are large, and the plan each account is on follows from its size rather than being sprinkled at random. Seats fit the plan, MRR is exactly seats times the plan's price, and nobody licenses more seats than they have employees. Support tickets resolve faster as priority rises.

TableRowsColumns
accounts1,200account_id, company_name, industry, country, employee_count, signup_date, plan
users21,884user_id, account_id, full_name, role, emailwork emails on the company's own domain
subscriptions1,200subscription_id, account_id, seats, status, mrr, started_on, ended_on
invoices14,500invoice_id, account_id, invoice_date, amount, status
support_tickets5,600ticket_id, account_id, opened_at, priority, category, satisfaction_score, resolution_hours, resolved_at

What was verified

  • 0 orphaned foreign keys, exactly one subscription per account
  • 0 invoices or tickets dated before the account existed
  • mrr equals seats times the plan's seat price, to the cent, for all 1,200
  • 0 accounts licensing more seats than they have employees
  • Seats rise with plan: 5, 23, 69, 189 median for Starter to Enterprise
  • Churned subscriptions all carry an end date; active ones never do
  • Median resolution: 4h urgent, 12h high, 34h normal, 77h low
  • 8.4% of tickets are still open, and none of those carry a satisfaction score
  • All 21,884 user emails are unique

Questions worth asking it

  • Does support load predict churn?
  • What is net revenue retention by plan tier?
  • Which industry has the worst satisfaction scores?
  • How does seat utilisation vary between Starter and Enterprise?

Regenerate it yourself with seed 20260722 for byte-identical output, or change the story and get a dataset nobody else has.

Need one shaped like your data?

These are fixed samples. If you need different tables, different volumes, or numbers that hit a specific target, describe it and generate it in the browser. No account needed to try.

Fully synthetic. No real person, company, or transaction is represented, and no production data was read to make these. Released under CC0: use them in portfolios, courses, demos, or products without attribution.