A preview only shows shape
A sample names its method
Aggregates default to the full set
LIMIT is not a random sample
Answers show row count and watermark
A sample can show structure, not the population
If the first 20 order rows have a high value per order, that is not the company average. A preview confirms columns, grain, and that rows exist. Averages, sums, shares, and “worst ten stores” must aggregate over an agreed population. When a tool sends only the preview to the model, the model can still state an average in a confident voice. That number describes those rows alone.
Enterprise queries should default to a full aggregate, or to a declared aggregate table. Return sample rows only when the user asks to see examples, and title them as examples. Kimball’s grain rule applies: decide what one row is before averaging those rows.
Preview, LIMIT, and a random sample are different
A preview is often the first rows in storage or primary-key order. New stores, test orders, or one partition show up first. PostgreSQL documents that LIMIT without a unique ORDER BY yields an unpredictable subset, and that LIMIT can change the plan. The average of LIMIT 100 is not excused by saying the sample is large.
A random sample is a third path. If the engine uses a mechanism such as TABLESAMPLE, the answer must name the method, the fraction or row count, and that the figure is an estimate. “I looked at some rows” is not a sample. A sum of sample amounts is not the population sum. Totals, counts, and figures that must conserve should not be extrapolated unless the business accepts an interval.
A column profile is not this query
Some systems precompute common values, null rates, and approximate cardinalities. Those profiles help ask whether a status column contains CLOSED. They do not answer “sales of CLOSED orders last week.” The profile’s watermark and filters rarely match the question. If the model treats frequent values as this week’s mix, it drops the long tail or promotes a historical mode into a current structure.
Use profiles for clarification and column choice. Take numbers from an aggregate with this question’s filters, and return both row counts. The NIST AI Risk Management Framework treats measurement as part of governance. A measurable failure is a total that the audited query cannot reproduce, or a query that clearly includes a preview limit.
Two paths, and when each applies
Path one aggregates the full authorized set. It fits operating numbers, reconciliation, and ranks. It costs more, and the answer can say “based on 128440 visible rows.” Path two is a labeled sample. It fits “what do rows look like” and early estimates of a distribution. The sample chart must not land in the daily operating pack, and the export must carry the sample mark.
Do not splice the paths in one sentence. Do not pair a full-population sales total with an average discount computed on the sample. If a pre-aggregate is used, state its grain, refresh time, and possible gap versus detail. Do not pretend it was a fresh full scan.
Counterexamples and acceptance
Counterexamples: an average from the first 50 rows; LIMIT 100 to find problem stores, with membership changing as the plan changes; a profile’s frequent channel described as this month’s main channel; sample amounts scaled up into a monthly forecast. Acceptance checks the aggregate scope, whether LIMIT or TABLESAMPLE is present, the row count, and the watermark. With preview access removed, an operating question should still aggregate fully or refuse.
The follow-up “drop those rows and average again” must filter the same population. It must not average the 20 rows sitting in the model context. AskTable.ai can be a candidate entry point. This article does not show that the current release already blocks previews and profiles from standing in for aggregates. Verify it on a table whose full sum is known.
Public references
Ready to help your team start?
Talk through a real scenario and see how AskTable.ai can fit your business.
Book a demo