AI-powered
podcast player
Listen to all your favourite podcasts with AI-powered features
How to Measure Test Coverage in a Software Project
With data, you can get the coverage metrics for your code, but it's not really straightforward how you would understand what that coverage metric looks like for your overall data. One of the challenges is that SQL isn't inherently composable in a sense and we can't run or test each column separately from the definition of the entire table. What I've seen being really effective in terms of setting standards for testing within teams is starting with really basic assumptions.
Data engineering is all about building workflows, pipelines, systems, and interfaces to provide stable and reliable data. Your data can be stable and wrong, but then it isn't reliable. Confidence in your data is achieved through constant validation and testing. Datafold has invested a lot of time into integrating with the workflow of dbt projects to add early verification that the changes you are making are correct. In this episode Gleb Mezhanskiy shares some valuable advice and insights into how you can build reliable and well-tested data assets with dbt and data-diff.
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Special Guest: Gleb Mezhanskiy.
Sponsored By:
Listen to all your favourite podcasts with AI-powered features
Listen to the best highlights from the podcasts you love and dive into the full episode
Hear something you like? Tap your headphones to save it with AI-generated key takeaways
Send highlights to Twitter, WhatsApp or export them to Notion, Readwise & more
Listen to all your favourite podcasts with AI-powered features
Listen to the best highlights from the podcasts you love and dive into the full episode