Who Chases the Data Quality Fix
Data teams can define the rules, build the gates and publish the dashboards, and the fixes still stall. What changed when the chasing moved to the owner of the numbers.
Numbers the business can see are wrong
In a lot of businesses, dashboards are already in place and shared with business users. Even in the cases where those dashboards are shared with users who own the targets and should act on them (covered in the first article of this series), the business users get data that is wrong in ways they can see: revenue not reconciled against their numbers, or stores or sites missing or wrongly named.
They have no clear way to quantify those data quality issues and size them to help them in their discussions with the data teams and upstream system owners.
When the business requests data quality improvements from upstream system owners, I saw those requests get added to the bottom of a backlog of what is judged as “higher priority” work, until they get buried and never looked at, and with no path to fixing the data and improving on it in the short to medium term.
These data quality issues cause the business users to lose faith in the numbers and go back to extracting data and using spreadsheets to help them in their day-to-day decisions. In several engagements, businesses were running multiple versions of the same number, data reconciliation issues happened all the time and user mistakes were induced by manual work.
The data team’s problem
In my experience, I saw different methods to tackle the data quality problem. In one organisation, we had a dedicated lead for data quality and governance. The lead had regular meetings with IT system owners and business leaders to discuss the downstream data issues and ways to fix them.
In those meetings, the data governance lead was going through reports of issues that they collected with examples and what needed addressing. The audience agreed about the need to fix those issues, but they rarely got prioritised and a couple of years down the line the same data quality problems were still present.
In one of the first engagements I led, we tried to add a more self-service approach to data quality as part of a wider platform and analytics delivery.
We worked with the product owner and the data team to define the business rules for the data we were going to use in the analytics, applied them as part of the data ingest and transformation pipelines with the outcomes as gates to stop bad data beyond a threshold reaching the business users, and created data quality dashboards that business leaders and IT system owners could access on their own to track the issues.
The approach used better tooling, which enabled faster feedback and allowed direct access to the people who could act on improving the data quality. The whole process was automated end to end, with people accessing the data quality dashboards that showed metrics about the data quality level and issues with exact rules that were breached, but the downstream fixes lagged.
The data team had to push hard with system owners to include any application fix. While the pushing from the data team was happening, some fixes were implemented, but as soon as that stopped, so did the fixes. We implemented a similar process on several other engagements, and the same outcome happened, which prompted us to start thinking about a different approach.
This is what I call a more traditional approach to data quality management, where it is led by the data team, from defining what the rules are, what fields those rules should apply to and owning the discussions with the business leaders and system owners.
When the owner started chasing
After facing the data quality challenges of the traditional approach, I started thinking about a different approach to our data quality process, where it could be owned by the business rather than by the data team, and to see if the data quality fixes would be better prioritised.
This coincided with a project where my team wanted to use machine learning, but our finding was that most of the features were incomplete and the overall data quality level was low. We then shifted our priority to focus on helping the organisation improve their data quality to be able to do analytics and machine learning on data they trust.
We worked with our client to define the domain to start with for the new data quality process, which held the main data the business owns, and we created a process to follow to improve its quality. This process started with identifying the business owner of the domain.
We then worked with the business owners and their teams to identify the important elements of data from that domain, where without them the business cannot operate or take decisions, such as activity dates, monetary values, categories, names.
This step was done based on survey forms to collect the information and validation workshops to finalise and agree the definitions, the owner of each data element, the business rules that apply for each one (e.g. name should not be empty, sales amount should be positive), and the target metric it impacts if that value is not correct (e.g. sales amount and related sales fields impact revenue, start date of an activity impacts time to close a task).
We ended up with a catalogue of important data elements and an information card for each one of them. Following the completion of the business catalogue, we worked with the source system owners to define the source field of each data element to the system and field where it originated from.
After completing the catalogue of important data elements with all the business definitions, rules, and lineage tracking to the system of record, the implementation of the data quality rules started and a dashboard was created. The dashboard showed the target metrics value of the impacted data records as the headline number, which was the business KPI in that domain (e.g. the revenue of the business or similar), rather than a compliance metric. The dashboard also included a data quality score and different filters and categories to help identify which part of the business needed to take action.
For the organisation we implemented this process for, I noticed it took the business owner only a few weeks after completing this process and seeing the dashboard to start having discussions with the system owners and push for data improvements.
Data quality moves only when the failures are priced in targets that the business owns and cares about. This new process changes the dynamic from having discussions with the system owners being owned by the data team to being owned by the business owner. For our example, it is a bit early to see the figures move, but the behaviour has already changed and it is going in the right direction.
Where to start
The first decision in this new process is what the failure costs and who owns those numbers.
Going into this exercise comes with an overhead for the people whose targets carry the failure numbers; business owners will have to dedicate time to this work and lean on their teams alongside them to define the data elements, what targets they impact, what failure means for the business and what good quality data is for those pieces of data.
In the engagement we discussed, the chasing moved from the data team to the business itself in weeks. If that move does not happen, it is worth reviewing the target metrics and the failure definition.
When the chasing moves, the trust in the data will start to be earned.