A CTO should treat data quality as a property of a decision path, not a score that a platform displays. A fulfillment leader deciding whether to expedite an order needs the right order, a usable status, and a freshness window that matches the shipping cut-off. A marketing aggregate that is one day late may be inconvenient; a credit exposure aggregate that silently omits a source may be unsafe. Start by listing the decisions that change money, customers, compliance, or operational capacity. For each one, name the records, transformations, people, and allowable delay that make the answer fit for use. This framing makes investment choices much clearer than a generic request to clean the data.
Define the decision before expanding data quality
The first boundary for data quality is the decision contract: who uses the result, what action they can take, when they need it, and what error is unacceptable. Turn that statement into a short review artifact with an accountable business owner and a technical owner. It should state the population, time basis, authoritative source, material exclusions, and a route for exceptions. This prevents a broad platform initiative from claiming success because it produced data, while the intended reader still relies on a spreadsheet or private interpretation. A narrow, repeated decision is the best starting point because it forces the team to make terms and handoffs concrete.
- Name the operator or leader who will change an outcome after seeing data quality.
- Describe the population and time rule in plain language, including exclusions.
- Identify the source or record that is authoritative when systems disagree.
- Set a freshness or review window that matches the action rather than a generic technical target.
- Write the fallback and escalation path for missing, contradictory, or restricted data.
Make data quality evidence inspectable
Write a promise at the grain people use. A customer master may require a durable identifier and a defined matching policy; an order fact may require one row per commercial order and explicit treatment of cancellations. Measure completeness, validity, uniqueness, consistency, timeliness, and accuracy only where each dimension has a stated consequence. Accuracy is often established by reconciliation to an accountable source or by a sampled business review, not inferred from a null-rate chart. The dbt guidance on data tests is useful here because it treats a test as an assertion that returns the records disproving it. Retain those failures long enough to find the source of a recurring defect.
| Design area | Decision to make | Evidence to keep |
|---|---|---|
| Decision input | Quality promise | Useful evidence |
| Dispatch queue | Every eligible order appears once before cut-off | Source-to-queue count and duplicate exceptions |
| Revenue report | Amounts reconcile to the approved ledger scope | Control total, period rule, and variance review |
| Customer outreach | Consent and contact status are current | Effective date, source authority, and suppression failures |
Build an operating path for data quality
Build quality controls at several points rather than hoping a final dashboard catches everything. Check incoming volume and schema at ingestion, keys and relationships during transformation, and control totals or business rules before publication. A late upstream file should surface as a stale result, not be replaced silently with an older feed. Put the quality status beside the measure that depends on it, with an owner and a route for investigation. This is especially important when batch retries, backfills, or manual corrections can make a technically completed job look current while its business period is incomplete.

Set controls and responses for data quality
Controls should test a declared promise and lead to a known response. For data quality, combine preventive controls, such as controlled schemas or access roles, with detective controls, such as reconciliation, freshness checks, and review of unexpected distributions. Do not make every deviation an incident; define materiality so teams can separate a correctable record from a decision-threatening condition. Each alert or review should identify the owner, affected scope, evidence available, containment choice, and communication expectation. The result is a service that can explain its limitations under pressure, not just a successful scheduled job.
| Control moment | Question | Expected response |
|---|---|---|
| Signal | Likely interpretation | First response |
| Freshness breach | Input is late or pipeline is stalled | Qualify the output, inspect arrival and job status |
| Relationship test failure | Key or reference rule changed | Quarantine affected records and trace the producer |
| Control-total variance | Scope, mapping, or period rule differs | Reconcile to accountable records before publication |
Work through a real data quality case
Consider a marketplace whose returns report suddenly improves after a product-system release. The useful response is not to average the new number with the old one. The team compares event counts by channel, checks whether return reasons changed, traces the transformation, and asks the product owner whether the new workflow emitted the same business event. If the definition changed, version the metric and restate only the period the policy supports. That incident creates a durable control: event-volume monitoring, a contract review before the next release, and a documented interpretation for finance and operations.
Govern change and access in data quality
Do not centralize every remediation decision in a data office. Source teams own the correctness of captured facts; domain owners approve meaning and acceptable exceptions; platform teams own reliable transport and observability; analytics teams own published transformations. A CTO should require an escalation route for a broken promise, a recovery target, and evidence that a correction reached every dependent output. The NIST data governance profile is a useful reminder that governance combines roles, policies, and practices rather than a catalog alone.
Measure whether data quality improves the work
Measure data quality through the quality of the decision path, not implementation activity alone. Useful signals include time from a material signal to a documented response, recurring disputes over a definition, percentage of decisions supported by current evidence, unresolved exceptions, and the number of parallel workarounds. Compare these with a baseline, then ask users to explain a representative result and what they would do if its main input were delayed. A higher dashboard view count or a larger catalog may be encouraging, but neither proves that decisions became more reliable. Revisit the measure when the workflow, source system, or ownership model changes.
Run the first 90 days of data quality deliberately
In the first month, choose one high-value workflow and establish its baseline: current preparation time, exception rate, decision delay, and the manual reconciliation that people perform today. In the second month, release the smallest complete data quality path to the people who already do that work. Include source status, an owner, a drill route, and a log for disputed cases; do not add broad self-service until these basics survive ordinary use. In the third month, review a sample of normal decisions, difficult exceptions, and a controlled failure such as a late input or a definition change. Record what the team learned, remove a workaround only after the replacement is reliable, and decide whether the same pattern is ready for a second domain. This sequence makes investment visible without rewarding superficial rollout activity.
Review the data quality operating system
A quarterly review keeps data quality aligned with the work rather than the original project plan. Bring together the business owner, source owner, technical operator, and a regular reader. Examine the most consequential incident, the most common reader question, meaningful changes to source scope or policy, access exceptions, and measures that no longer lead to action. Verify that contact details and runbooks still work, that failed checks retain enough evidence for investigation, and that historical comparisons carry the right definition label. Decide explicitly whether to tighten a promise, accept a bounded limitation, automate a repeated check, or retire a stale output. The review should leave a short record of decisions and owners, so the next change starts with context instead of rediscovery.
Make the next data quality decision easier
Use the review to remove friction for the next person who needs data quality. Add a concise definition where a reader hesitated, preserve a representative failing record where an incident was difficult to reproduce, and put the owner or escalation contact beside the output that needs it. When a workaround has become routine, decide whether it represents a missing product feature, an unavoidable control, or a path that should be retired. This small discipline prevents institutional knowledge from living only in chat messages and meeting memory. It also makes scale more realistic: a new team can adopt an established decision pattern with its boundaries, evidence, and response practice already visible.
Key takeaways for data quality
- Start data quality with a real decision, named owner, and explicit time requirement.
- Make source authority, definitions, scope, and limitations visible near the result.
- Test declared promises at the source, transformation, and publication points.
- Treat exceptions, late data, and semantic changes as design cases rather than edge cases.
- Use incidents and reader questions to improve the next release instead of accumulating undocumented workarounds.
Frequently asked questions about data quality
Who owns data quality? Ownership is shared but not vague: a business owner approves the decision meaning, source owners protect captured facts, and technical owners operate the path and controls. How broad should a first release be? Make it narrow enough to test in one working cadence, but complete enough to include authority, quality checks, access, and an exception route. When should a definition change? Change it when the business meaning genuinely changes; version the rule, compare results where practical, and tell affected readers the effective date. What should happen when data is late? Show the status, follow the agreed fallback or hold rule, and investigate the cause instead of presenting a silently stale answer.
Conclusion: make data quality a maintained decision capability
CTOs get the greatest return from data quality when they build it as a maintained capability: a bounded decision, inspectable evidence, explicit controls, a response owner, and a learning loop. Begin with the path that is already causing friction, document its promises, and prove the workflow with ordinary and difficult cases. Then expand only after the team can explain a result, recover from a known failure, and show that the decision improved. That approach keeps technical ambition connected to the people, records, and consequences that make the data worth trusting.