> ## Documentation Index
> Fetch the complete documentation index at: https://docs.textql.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Data Quality & Validation

> A workshop of short, hands-on modules for data engineers and analytics engineers who need to trust — and prove — answer accuracy. The thesis throughout: accuracy is checked, not…

A workshop of short, hands-on modules for **data engineers and analytics engineers** who need to trust — and *prove* — answer accuracy. The thesis throughout: **accuracy is checked, not asserted.** Golden datasets pinned to numbers the org already trusts, eval cases that exercise edge conditions, drift detection on a schedule, reconciliation workflows for when systems disagree, and a validation agent that runs the moment data lands.

Total time: **about 2 hours.**

Builds on [**Build Your Ontology, End to End**](../build-your-ontology/) and [**Ontology Operations**](../ontology-operations/) — this workshop assumes governed definitions exist and doesn't re-teach the build; it teaches the proof.

## Who this is for

* Data/analytics engineers accountable for "is this number right?"- Ontology owners who need tests, not vibes, behind their definitions- Teams burned by a silent upstream change that shipped wrong numbers for a week- Anyone preparing for an audit, a board cycle, or a regulator

## What you'll be able to do

<table><tr><th>After module</th><th>You can...</th></tr><tr><td>0</td><td>Map every way an answer goes wrong to the defense that catches it</td></tr><tr><td>1</td><td>Build golden queries pinned to numbers the org already trusts</td></tr><tr><td>2</td><td>Generate and curate eval cases that exercise definitions and edge cases</td></tr><tr><td>3</td><td>Run drift detection on a schedule, distinguishing data drift from definition drift</td></tr><tr><td>4</td><td>Reconcile two disagreeing systems to the exact fork, and record the resolution</td></tr><tr><td>5</td><td>Deploy a data-quality agent triggered by ETL completion</td></tr><tr><td>6</td><td>Write the quality runbook: owners, cadences, and what a red alert triggers</td></tr></table>

## Workshop modules

<table><tr><th>Module</th><th>Time</th></tr><tr><td>0 · The Trust Stack</td><td>10 min</td></tr><tr><td>1 · Golden Datasets</td><td>20 min</td></tr><tr><td>2 · Validation Sets & Eval Cases</td><td>20 min</td></tr><tr><td>3 · Drift Detection on a Schedule</td><td>20 min</td></tr><tr><td>4 · Reconciliation Workflows</td><td>20 min</td></tr><tr><td>5 · The Data-Quality Feed Agent</td><td>15 min</td></tr><tr><td>6 · The Quality Runbook</td><td>15 min</td></tr></table>

## How to use this workshop

* Prerequisites: a connected warehouse and an existing ontology with governed metrics (the build workshops above if not).- Work **in order** — golden datasets (Module 1) are the foundation every later module schedules, extends, or operationalizes.- Swap `[bracketed]` placeholders for your metrics, tables, and trusted sources.

<Accordion title="For facilitators — running this as a guided session">
  **T-minus-1-day pre-flight:** run the workshop's key prompts end-to-end in the session workspace — confirm logins, connectors, and the features this workshop touches are enabled for every attendee; have a fallback demo workspace ready in case a customer connector fails live. **Pacing:** when a long prompt is running (2–4 min), fire it first, then discuss the concept while Ana works — never watch a spinner in silence. If behind schedule, cut sections marked optional, never the checkpoints. **Group sessions:** attendees drive, you narrate; collect every miss or wrong answer in a shared doc — that list is the ontology backlog. With 10+ attendees on one warehouse, stagger the heavy prompts.
</Accordion>

<Note>
  **🤖 Prefer to have Ana run this workshop?** Open a thread with your data connected and paste the prompt below — Ana walks you through it on your own data, one step at a time. You run each prompt; she coaches. *(Air-gapped / VPC tenants where Ana can't reach the web: download [`ana-runner.md`](https://github.com/TextQLLabs/workshops/blob/main/data-quality/ana-runner.md) and paste its module list instead.)*
</Note>

```text Run with Ana theme={null}
Hey Ana — facilitate the "Data Quality" workshop with me in this thread, on the data connected here. Pull the steps from https://textqllabs.github.io/workshops/data-quality/ and run it interactively: look at my data first (2–3 lines), then go ONE module at a time. For each module, give ME the prompt to copy and run as my next message, tell me what to look for, then wait and coach me on my result. Start with what you see, then Module 0.
```

[**▶ Open in Ana — prompt loaded**](https://app.textql.com/chat/new?prefill=Hey%20Ana%20%E2%80%94%20facilitate%20the%20%22Data%20Quality%22%20workshop%20with%20me%20in%20this%20thread%2C%20on%20the%20data%20connected%20here.%20Pull%20the%20steps%20from%20https%3A//textqllabs.github.io/workshops/data-quality/%20and%20run%20it%20interactively%3A%20look%20at%20my%20data%20first%20%282%E2%80%933%20lines%29%2C%20then%20go%20ONE%20module%20at%20a%20time.%20For%20each%20module%2C%20give%20ME%20the%20prompt%20to%20copy%20and%20run%20as%20my%20next%20message%2C%20tell%20me%20what%20to%20look%20for%2C%20then%20wait%20and%20coach%20me%20on%20my%20result.%20Start%20with%20what%20you%20see%2C%20then%20Module%200.)

Ana fetches the steps from the link — no copy/paste of the workshop needed. (Air-gapped / VPC tenants where Ana can't reach the web: paste the module list from this workshop's [ana-runner.md](https://github.com/TextQLLabs/workshops/blob/main/data-quality/ana-runner.md) instead.)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.