Question: real recurring phrasing
Identity: role and data scope
Expectation: metric, filter, time
Execution: query, result, explanation
Decision: exact, tolerance, review
Test business meaning, not merely executable SQL
A model upgrade, renamed field, revised metric, or edited knowledge entry can change the interpretation of the same question.
Store question, identity, data snapshot, expected metric, filters, time window, and tolerance as the test contract.
Build a layered set from real demand
Include canonical wording, abbreviations, typos, vague dates, follow-ups, clarification cases, refusals, and insufficient-data cases.
Tier tests by domain and risk so critical finance or access cases run on every change while long-tail phrasing is sampled.
Score the answer in components
Compare metric, dimension, filter, join, time, access, value, chart, and explanation separately. Prose can vary while facts and constraints remain fixed.
Use explicit tolerances for floating point, late data, and live feeds, with a recorded baseline version.
Trace failures to the changed layer
Retain semantic parse, plan, query, result, and final response to distinguish model drift from data or policy change.
Block high-risk regressions and route lower-risk differences to review through versioned, staged release.
Exercise one real upgrade
Run the complete suite through a model or metric upgrade and verify attribution, approval, exceptions, rollback, and re-release.
AskTable.ai can be evaluated for governed query and semantics; automated comparison, thresholds, and release gates require current product confirmation.
Public references
Ready to help your team start?
Talk through a real scenario and see how AskTable.ai can fit your business.
Book a demo