Why data quality depends on meaning, evidence, and process
This is the fourth article in a five-part series on implementing HACT’s UK Housing Data Standards.
The earlier articles considered why organisations need to understand their own data, what a useful data dictionary should contain, and what the dictionary-building process can expose about definitions, ownership, and hidden business rules.
This article focuses on data quality.
A field can be complete, correctly formatted, and populated with a permitted value while still being inaccurate, out of date, or unsuitable for the purpose for which it is being used.
Understanding that difference is essential when assessing whether internal data can be aligned reliably with a common standard.
A valid value can still be wrong
Consider the following records:
- A postcode follows the correct format but belongs to another property
- A property type uses an active code but describes the wrong characteristic
- A property status is populated but has not been updated since the process changed
- An ownership code is permitted but does not reflect the relevant legal position
- A mandatory field contains a default value whenever the true answer is unknown
- A valid date represents a different event from the one assumed by a report
- A component is linked to a valid property, but not the property where it is installed
These records may pass standard completeness and validity checks.
They may still fail an accuracy or fitness-for-purpose test.
This distinction matters because organisations often begin data-quality work by measuring what can be tested most easily. They count missing values, invalid formats, duplicates, and codes outside an approved list.
Those checks are necessary, but they do not prove that the information represents the real-world condition assumed by the organisation.
Technical validity is not accuracy
Technical checks can often establish whether a value:
- Is present
- Follows an expected format
- Belongs to a permitted code list
- Falls within an allowed range
- Is unique where required
- Has a related record in another table
Accuracy normally requires more context.

A value can be present, correctly formatted, and drawn from an approved code list while still being inaccurate or unsuitable for its intended use. Technical validity is necessary, but accuracy also depends on meaning, evidence, timeliness, and contex
The organisation needs to understand what the value is intended to represent, where it should come from, what event creates or changes it, how frequently it should be reviewed, and which evidence can confirm it.
That is where the dictionary becomes important.
For each critical data item, it can identify:
- The intended meaning
- The expected source
- The process that creates it
- The required frequency of update
- The authoritative evidence
- The relationships that should hold
- Legitimate exceptions
- The owner responsible for deciding whether the value is acceptable
Without this context, an organisation may be able to report that a field is 99.9 per cent complete while remaining unable to say whether the values are correct.
The definition determines the quality rule
Consider an Occupied property status.
If Occupied means that an active tenancy exists, a quality rule might compare the property status with tenancy records.
If it means that a person is physically resident, the same rule is insufficient. A tenancy may continue during a temporary absence, a decant, an abandonment investigation, or a delayed termination.
If Occupied means that the property is unavailable for letting, the rule may need to consider a wider range of operational statuses and restrictions.
The field name has remained the same, but each definition produces a different quality rule.

The meaning assigned to a data item determines the evidence, validation rule, and exception process needed to assess its quality. The same status can require entirely different checks depending on what it is intended to represent.
This principle applies throughout housing data.
Before checking whether Ownership is accurate, the organisation must define which form of ownership is being assessed.
Before validating a Block hierarchy, it must decide whether the relationship is physical, administrative, financial, or compliance-based.
Before testing whether Property Type is correct, it must establish whether the field describes dwelling form, building use, tenure, construction, or operational purpose.
The definition determines the quality rule.
This is why data profiling cannot replace semantic understanding. An analysis can be technically correct while answering the wrong business question.
Quality has several dimensions
Data-quality discussions often focus on whether a record is complete or valid, but different uses require different dimensions to be considered.
A field might be:
- Complete, because every record contains a value
- Valid, because every value appears in the permitted code list
- Consistent, because related systems broadly agree
- Timely, because changes are recorded within an agreed period
- Accurate, because the value is supported by appropriate evidence
- Fit for purpose because it can support the intended decision or process
A field can perform well in one dimension and poorly in another.
For example, a mandatory Property Type field may be 100 per cent complete and use only valid codes. If the code list combines several concepts, the field may still be unsuitable for standards mapping or compliance population identification.
Similarly, two systems may agree because one copies data from the other. That demonstrates consistency, but not necessarily accuracy.
Quality measures therefore need to be tied to defined meaning, evidence, process, and intended use.
Quality problems may originate in the process
The dictionary also helps move the conversation away from treating poor data as a problem caused only by individual users.
A value may be wrong because:
- The system does not offer the correct option
- The information is not available when the record is created
- Two processes update the same field
- An interface overwrites a more recent value
- A mandatory field encourages the use of defaults
- Responsibility for review is unclear
- Data is copied from a source that is already outdated
- A business rule changed without the system being updated
Cleaning the records may improve the immediate dataset, but the problem will return if the process remains unchanged.
If a mandatory field is populated with a default because users do not know the answer at the point of entry, removing the defaults does not solve the underlying problem. The organisation may need to change when the information is collected, provide a valid Unknown status, introduce a review step, or redesign the system workflow.
If an interface repeatedly overwrites newer values, retraining users will have little effect.
By connecting quality issues to definitions, processes, systems, and ownership, the dictionary helps determine whether the required response is cleansing, redesign, training, system configuration, integration, or governance.
Some quality cannot yet be measured
A useful dictionary should allow the organisation to record that a quality expectation has not yet been established.
Possible findings include:
- Rule not defined
- Evidence source not agreed
- Threshold not approved
- Timeliness requirement unknown
- Exception process not documented
- Quality owner not assigned
- Authoritative source disputed
These are not desirable long-term positions, but they are legitimate findings.
The alternative is to invent a rule because the template requires one, or to adopt an existing definition without confirming that it reflects current business practice.
That creates false confidence.
The dictionary will not fix data quality by itself. It will show where quality cannot currently be defined, measured, evidenced, explained, or assigned to anyone.
That visibility is the starting point for improvement.
Unknown is a legitimate answer
There can be pressure during dictionary and standards work to complete every field, nominate an owner, define a quality rule, and select a corresponding standard attribute.
A fully populated spreadsheet can look reassuring, but it may conceal uncertainty rather than resolve it.
Where the evidence is incomplete, values such as Unknown, Under Review, or Not Yet Agreed may be the most responsible answer.
The important requirement is that uncertainty creates an action.
The organisation may need to:
- Analyse the actual data
- Trace an interface or transformation
- Review system configuration
- Consult operational and subject specialists
- Examine historic documentation
- Agree an evidence source
- Assign an owner
- Decide whether the field should be redesigned or retired
An explicit unknown can be investigated and resolved.
A confident but unsupported answer may remain unchallenged until it creates a reporting, integration, migration, or compliance problem.
Reliable standards mapping therefore depends on more than establishing that a similar field exists. The organisation must also know whether the data is sufficiently defined, evidenced, controlled, and fit for the intended purpose.
Once that understanding is in place, the organisation can begin assessing the nature of its alignment with the standards and deciding what action is required.
Part 5 will consider how to classify that alignment, choose a practical starting scope, and move from documentation into implementation.
About HACT
HACT is the charity of the social housing sector, supporting innovation, collaboration, and insight across housing. Since 2018, it has worked with OSCRE and organisations across the sector to develop the UK Housing Data Standards, which provide the common framework at the centre of this series. More recently, HACT has established UK HIVE to support sector-wide collaboration, research, development, and implementation.
I would like to acknowledge HACT, OSCRE, and the wider sector contributors whose work has made the standards possible. This series is an independent discussion of the practical questions that can arise when applying them to existing housing systems, processes, and data.

